# Gaia Architecture Comparison > Comparación lado a lado de ruflo frente a HAL y otros harnesses de GAIA: huecos de capacidad, decisiones de diseño y hoja de ruta de mejoras. Fuente: https://skillsagentes.com/skills/ruvnet/ruflo/gaia-architecture-comparison Markdown: https://skillsagentes.com/skills/ruvnet/ruflo/gaia-architecture-comparison.md Repositorio: https://github.com/ruvnet/ruflo Autor: ruvnet Licencia: MIT Actualizado: el mes pasado Coste de contexto: 32 tok instalada, 1.3k tok al activarse, 1.3k tok con todos los archivos del bundle Bundle: 1 archivo, 5 KB Permisos que pide: bash read mcp__plugin_ruflo-core_ruflo__memory_search mcp__plugin_ruflo-core_ruflo__memory_store ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent claude-code # Cursor npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent cursor # Codex npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent codex # Gemini CLI npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent gemini # Windsurf npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent windsurf # Cline npx -y skills add ruvnet/ruflo --skill gaia-architecture-comparison --agent cline ``` ## Qué hace - Compara la arquitectura del harness de benchmark GAIA de ruflo con la referencia HAL de Princeton y otros harnesses open-source - Detalla diferencias clave (número de preguntas, búsqueda web, ejecución de código, OCR, memoria) en una tabla - Prioriza mejoras en una hoja de ruta con impacto y esfuerzo estimados ## Cuándo usarla - Planificas la siguiente iteración del trabajo en GAIA - Evalúas qué cambio arquitectónico tiene mayor ROI en la tasa de aciertos - Incorporas a un nuevo colaborador al código del benchmark ## Qué la activa - "Compara la arquitectura de nuestro harness GAIA con la de HAL" - "¿Qué cambio nos daría más ROI en el pass rate de GAIA?" ## Antes de instalar - Necesita en el PATH: npx ## Archivos - SKILL.md — 5 KB ## SKILL.md Reproducido tal cual desde ruvnet/ruflo bajo MIT. Esta sección es el documento original y está en inglés. # GAIA Architecture Comparison Skill Compare ruflo's GAIA benchmark harness against the Princeton HAL reference implementation and other open-source harnesses to understand capability gaps and prioritize improvements. ## When to use - Planning the next iteration of GAIA work - Evaluating which architectural change has the highest pass-rate ROI - Onboarding a new contributor to the benchmark codebase ## Architecture overview ### ruflo harness (current) ``` gaia-bench run └─ gaia-loader.ts — HF dataset download + cache └─ gaia-agent.ts — multi-turn Anthropic Messages loop └─ gaia-tools/ — web_search, file_read, web_browse, image_describe, python_exec └─ gaia-voting.ts — Track A self-consistency (N attempts → majority vote) └─ gaia-hardness/ — Track Q difficulty predictor (ADR-136) └─ gaia-judge.ts — two-stage LLM-as-judge scorer ``` ### HAL reference (Princeton) HAL uses a similar loop but with: - OpenAI function calling as the tool interface - BrowserBase / Playwright for real browser automation - Code interpreter sandbox (Jupyter kernel) - Larger token budget per turn (4096+) - Full 300-question evaluation set ### Key differences | Dimension | ruflo | HAL reference | Gap | |-----------|-------|--------------|-----| | Question count | 53 (partial L1) | 300 (full L1) | Use `--limit 165` for full L1 | | Web search | DuckDuckGo / Google CSE | BrowserBase live | Add Playwright or Browserless | | Code execution | python_exec stub | Real Jupyter kernel | Implement real sandbox | | Image OCR | image_describe (Gemini) | GPT-4V / Gemini | Functionally equivalent | | File handling | file_read | Full PDF/XLSX/ZIP parser | Expand file_read | | Self-consistency | voting.ts (Track A) | Not in reference | ruflo advantage | | Hardness routing | predictor.ts (Track Q) | Not in reference | ruflo advantage | | Memory | AgentDB HNSW | None | ruflo advantage | | Pass-rate L1 | ~20.8% (iter 23) | 74.6% (HAL Sonnet 4.5) | ~54 pp gap | ## Gap analysis ### Primary gaps (high impact) 1. **Real code execution** — many L2/L3 questions require running Python to compute a numerical answer. The current `python_exec` tool is a stub. Implementing a real sandbox (E2B, Pyodide, or subprocess) is the single highest-ROI change. 2. **Full question set** — running 53/300 L1 questions underestimates true pass-rate because the first 53 skew easier. Run `--limit 165` (full L1) for a comparable HAL score. 3. **Real browser** — `web_browse` currently fetches raw HTML. Replacing it with Playwright/Browserless for JavaScript-rendered pages would unlock many web navigation questions. ### Secondary gaps (medium impact) 4. **Structured file parsing** — PDF, XLSX, and ZIP attachments require dedicated parsers. `file_read` currently handles plain text and images only. 5. **Turn budget** — 12 turns may be insufficient for complex multi-step questions. HAL uses up to 20 turns for L3. 6. **System prompt tuning** — HAL's system prompt is more elaborate and explicitly instructs the model to use tools before answering. ### ruflo advantages 7. **Self-consistency voting** (Track A) — running N attempts per question and taking the majority answer reduces variance on borderline questions. HAL does not implement this. 8. **Hardness routing** (Track Q) — routing each question to an appropriate model and turn budget based on predicted difficulty. This reduces cost on easy questions while providing more resources for hard ones. 9. **AgentDB memory** — storing patterns across runs enables the agent to recall successful strategies for similar question types. ## Improvement roadmap | Priority | Change | Expected Lift | Effort | |----------|--------|--------------|--------| | P0 | Real python_exec sandbox (E2B) | +15-25 pp | High | | P0 | Full 165-Q L1 evaluation | Accurate baseline | Low | | P1 | Playwright-based web_browse | +5-10 pp | Medium | | P1 | PDF/XLSX file parser | +3-8 pp | Medium | | P2 | Increase max-turns to 20 for L2/L3 | +2-5 pp | Low | | P2 | System prompt tuning (iter 30 research) | +2-5 pp | Low | | P3 | Google Grounding via Gemini (iter 32) | +3-7 pp | Medium | | P3 | Multi-provider routing (Gemini Flash for cheap Q's) | Cost reduction | Medium | ## Loading context from past research ```bash npx @claude-flow/cli@latest memory search \ --namespace gaia-patterns \ --query "architecture comparison HAL benchmark" ``` ## Storing comparison findings ```bash npx @claude-flow/cli@latest memory store \ --namespace gaia-patterns \ --key "architecture-comparison-$(date +%Y%m%d)" \ --value "HAL gap: 54pp. Primary: python_exec stub. Secondary: browser, file parsing." ``` ## Dónde encaja - Categoría: [Investigación](https://skillsagentes.com/categorias/investigacion.md) — Investigación estructurada, búsqueda de fuentes y síntesis. - Creador: [ruvnet](https://skillsagentes.com/creators/ruvnet.md) — 275 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Harness Gepa](https://skillsagentes.com/skills/ruvnet/ruflo/harness-gepa.md): Inspecciona y audita genomas GEPA: carga y valida un genoma, renderiza el system prompt que compila, o clasifica los modos de fallo de una transcripción de ejecución. - [Deepseek Reason](https://skillsagentes.com/skills/ruvnet/ruflo/deepseek-reason.md): Completion en modo razonamiento contra deepseek-reasoner (R1) de DeepSeek. Devuelve el chain-of-thought por separado de la respuesta final. Lee DEEPSEEK_API_KEY y degrada si falta o la API no responde. - [Deepseek Chat](https://skillsagentes.com/skills/ruvnet/ruflo/deepseek-chat.md): Completion de un solo turno contra el modelo deepseek-chat de DeepSeek vía /v1/chat/completions. Lee DEEPSEEK_API_KEY y degrada con status:degraded si falta o la API no responde. Para tareas sin razonamiento. - [Adr Index](https://skillsagentes.com/skills/ruvnet/ruflo/adr-index.md): Construye o reconstruye el índice de ADRs y su grafo de dependencias ejecutando scripts/import.mjs, en vez de cientos de llamadas MCP. - [Agntcy Status](https://skillsagentes.com/skills/ruvnet/ruflo/agntcy-status.md): Muestra el estado de la integración AGNTCY/SLIM/CASA: si los paquetes están instalados, qué transporte está activo y si el enforcement de CASA está habilitado. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)