# Ai Research Explore > Slug de skill compatible con Rigor Explore para candidatos de investigación en deep learning potencialmente novedosos: exploración solo de candidatos sobre `current_research`, con comprensión auditable del repo y comparación justa. Fuente: https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/ai-research-explore Markdown: https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/ai-research-explore.md Repositorio: https://github.com/lllllllama/RigorPilot-Skills Autor: lllllllama Licencia: MIT Actualizado: hace 2 meses Coste de contexto: 153 tok instalada, 1.6k tok al activarse, 81.3k tok con todos los archivos del bundle Bundle: 34 archivos, 318 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent claude-code # Cursor npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent cursor # Codex npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent codex # Gemini CLI npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent gemini # Windsurf npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent windsurf # Cline npx -y skills add lllllllama/RigorPilot-Skills --skill ai-research-explore --agent cline ``` ## Qué hace - Ejecuta exploración candidate-only de investigación en deep learning sobre un anclaje `current_research` con idea gating, comparación justa y experimentos gobernados - En modo campaign, congela task, dataset, benchmark, evaluation source, SOTA reference y budget antes del trabajo de candidatos - Usa un ritmo de dos loops: uno externo para entender el repo y filtrar ideas, otro interno para cambios acotados y recolección de evidencia - Escribe resultados solo de candidatos en `analysis_outputs/`, `sources/` y `explore_outputs/`, incluyendo `SCIENTIFIC_CHANGELOG.md` y `COMPARABILITY_REPORT.md` - Prioriza candidatos antes de ejecutar por ganancia y costo esperados, y después por evidencia real ## Cuándo usarla - El investigador ya eligió task family, dataset, benchmark, método de evaluación y referencias SOTA, y quiere exploración solo de candidatos - Hay autorización explícita de exploración (trabajo candidate-only, branch o worktree aislado, sweep, varias variantes o ranking exploratorio) - Existe un contexto `current_research` durable como branch, commit, checkpoint, run record o modelo local ya entrenado ## Cuándo no - Reproducción confiable basada en README, búsqueda abierta de dirección, exploración estrecha solo de código o solo de ejecución, análisis pasivo del repo, afirmaciones de novedad verificadas, o experimentación implícita - Peticiones estrechas de solo código van a `explore-code`, solo de ejecución a `explore-run`, análisis pasivo a `analyze-project`, reproducción basada en README a `ai-research-reproduction` ## Qué la activa - "Explora variantes candidatas sobre mi current_research congelando dataset, benchmark y SOTA reference" - "Necesito un research_campaign para comparar candidatos frente a la baseline actual con evidencia auditable" - "Quiero explorar mejoras candidatas sobre este checkpoint entrenado con budget de cómputo fijo" ## Antes de instalar - Requiere una solicitud autorizada de exploración candidate-only más un anclaje `current_research` durable (branch, commit, checkpoint o run record). - makes network requests ## Archivos - SKILL.md — 6 KB - agents/openai.yaml — 423 B - references/ai-research-explore-policy.md — 2 KB - references/idea-evaluation-framework.md — 1 KB - references/research-campaign-spec.md — 9 KB - references/smoke-validation-policy.md — 787 B - references/source-mapping-policy.md — 615 B - references/sources-naming-policy.md — 895 B - scripts/lookup/__init__.py — 624 B - scripts/lookup/cache_store.py — 8 KB - scripts/lookup/inventory_writer.py — 3 KB - scripts/lookup/normalizers.py — 5 KB - scripts/lookup/providers/__init__.py — 457 B - scripts/lookup/providers/arxiv_provider.py — 3 KB - scripts/lookup/providers/base.py — 3 KB - scripts/lookup/providers/doi_provider.py — 3 KB - scripts/lookup/providers/github_provider.py — 3 KB - scripts/lookup/providers/optional_provider.py — 844 B - scripts/lookup/providers/url_provider.py — 2 KB - scripts/lookup/record_schema.py — 3 KB - scripts/lookup/repo_extractors.py — 3 KB - scripts/lookup/source_support.py — 4 KB - scripts/orchestrate_explore.py — 118 KB - scripts/passes/__init__.py — 909 B - scripts/passes/atomic_idea_decomposition.py — 13 KB - scripts/passes/candidate_idea_generation.py — 22 KB - scripts/passes/execution_feasibility.py — 19 KB - scripts/passes/idea_cards.py — 2 KB - scripts/passes/idea_ranking.py — 8 KB - scripts/passes/implementation_fidelity.py — 16 KB - scripts/passes/improvement_bank.py — 19 KB - scripts/passes/lookup_sources.py — 16 KB - scripts/passes/source_mapping.py — 19 KB - scripts/write_outputs.py — 830 B ## SKILL.md Reproducido tal cual desde lllllllama/RigorPilot-Skills bajo MIT. Esta sección es el documento original y está en inglés. # ai-research-explore ## Purpose Use this as the Rigor Explore compatible skill slug after the researcher explicitly authorizes candidate-only work on top of a durable `current_research` anchor. The installed slug remains `ai-research-explore` for compatibility. Rigor Explore is for meaningful and potentially novel deep learning research candidates while preserving scientific rigor, comparability, reproducibility, and auditable collaboration. Novelty and significance remain hypotheses before literature contrast, ablation evidence, and fair comparison. The skill does not promise autonomous discovery, global benchmark completeness, novelty proof, or trusted reproduction success. Start from the shared operating principles in `../ai-research-reproduction/references/agent-operating-principles.md`, then load `../ai-research-reproduction/references/research-rigor-principles.md` for research claims and `../ai-research-reproduction/references/deep-learning-experiment-principles.md` when experiment details affect comparability or reproducibility. ## Fit Use this skill only when the request has both: - Explicit exploration authorization such as candidate-only work, isolated branch or worktree, sweep, several variants, or exploratory ranking. - A durable `current_research` context such as a branch, commit, checkpoint, run record, or already-trained local model state. Keep narrow code-only requests on `explore-code`. Keep narrow run-only requests on `explore-run`. Keep passive repository analysis on `analyze-project`. Keep README-first reproduction on `ai-research-reproduction`. ## Research Rhythm Use a two-loop rhythm: - Outer loop: understand the repository, freeze task/dataset/evaluation/budget, preserve user ideas, map sources, gate ideas, and decide whether the next experiment is worth running. - Inner loop: make one bounded candidate change or run, smoke-check it, collect evidence, rank it against the current anchor, and either stop or return to the outer loop with the new evidence. This rhythm is a guide, not a rigid autonomous loop. Stop at explicit blockers, unclear scientific meaning, exhausted budget, missing anchor/evaluation, or a human checkpoint. ## Workflow 1. Confirm `current_research` and explicit explore-lane authorization. 2. Accept either legacy `variant_spec` or higher-level `research_campaign`. 3. In campaign mode, freeze the task, dataset, benchmark, evaluation source, SOTA reference, and budget before candidate work. 4. Build only the repo-understanding artifacts needed for the current campaign, usually through `analyze-project`. 5. Run bounded, cache-first source lookup when source support matters; prefer local curated literature such as Zotero if available, then seed sources, repo-local locators, public locators, or optional web lookup. Treat lookup as source resolution, not an open-ended literature search. 6. Preserve researcher-provided ideas, optionally add a small bounded set of single-variable seed ideas, and rank ideas with explicit gates and score breakdowns. 7. Prefer one clear candidate at a time. Use `explore-code` for bounded code adaptation and `explore-run` for short-cycle trials or sweeps. 8. Use `minimal-run-and-audit` or `run-train` only when the exploratory plan requires real execution evidence. 9. Write candidate-only outputs to `analysis_outputs/`, `sources/`, and `explore_outputs/` as appropriate; never present exploratory gains as trusted reproduction success. Include `SCIENTIFIC_CHANGELOG.md` and `COMPARABILITY_REPORT.md` for candidate scientific meaning and comparison boundaries. ## Ranking and Evidence - Before execution, prioritize candidates by expected gain, cost, success likelihood, patch surface, dependency drag, evaluation risk, and rollback ease. - After execution, rank by real evidence first: command status, observed metrics, artifacts, changed paths, smoke results, and reproducibility notes. - Keep researcher-provided `evaluation_source` and `sota_reference` frozen for the campaign; do not claim they are globally complete. - If the top ideas are too close or the implementation cannot be decomposed into auditable units, stop for a checkpoint instead of silently choosing. ## Campaign Inputs `research_campaign` is preferred for Rigor Explore campaigns, but it should stay minimal. The durable core is: - `current_research` - `task_family` - `dataset` - `benchmark` - `evaluation_source` - `sota_reference` - `compute_budget` Use `candidate_ideas`, `variant_spec`, `research_lookup`, `idea_policy`, `idea_generation`, `source_constraints`, `feasibility_policy`, `baseline_gate`, and `execution_policy` as optional guidance, not as fields the agent must fill for every campaign. See `references/research-campaign-spec.md` for the advanced schema and artifact expectations. ## Reference Loading - Load `references/ai-research-explore-policy.md` for lane safety and candidate semantics. - Load `references/research-campaign-spec.md` only when a campaign file is present or the user asks for Rigor Explore campaign governance. - Load `../ai-research-reproduction/references/explore-variant-spec.md` for run-level variant matrix details. - Load `../ai-research-reproduction/references/research-thinking-loop.md` before proposing or ranking candidate changes; it is the required greedy observe-ground-design-compare cycle. - Load `../ai-research-reproduction/references/research-rigor-principles.md` before making novelty, contribution, SOTA, or comparability statements. - Consult `~/.rigorpilot/PERSONAL_RIGOR.md` if present, under `../ai-research-reproduction/references/continuous-learning-policy.md` (advisory only; core wins). - Load `../ai-research-reproduction/references/deep-learning-experiment-principles.md` when training, evaluation, baseline, ablation, metric, checkpoint, or dataset details matter. - Use `scripts/orchestrate_explore.py` and `scripts/write_outputs.py` for the existing deterministic artifact workflow. ## Dónde encaja - Categoría: [Investigación](https://skillsagentes.com/categorias/investigacion.md) — Investigación estructurada, búsqueda de fuentes y síntesis. - Creador: [lllllllama](https://skillsagentes.com/creators/lllllllama.md) — 11 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Ai Research Reproduction](https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/ai-research-reproduction.md): Skill compatible con Rigor Reproduce para reproducción README-first de repos de deep learning: elige el objetivo mínimo confiable, coordina fases confiables y registra evidencia y desviaciones en `repro_outputs/`. - [Repo Intake And Plan](https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/repo-intake-and-plan.md): Ayudante de Rigor Intake para reproducción de repos de deep learning basada en README: escanea el repo, extrae comandos documentados y clasifica candidatos de inferencia, evaluación y entrenamiento. - [Minimal Run And Audit](https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/minimal-run-and-audit.md): Skill Rigor Run para reproducir repos de deep learning centrados en el README: captura evidencia de un smoke test, inferencia o evaluación documentada en repro_outputs/, con notas de parches si cambian archivos. - [Run Train](https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/run-train.md): Skill de Rigor Train para repositorios de investigación de deep learning: ejecuta de forma conservadora comandos de entrenamiento y registra evidencia estandarizada en train_outputs/. - [Explore Run](https://skillsagentes.com/skills/lllllllama/rigorpilot-skills/explore-run.md): Skill hoja de Rigor Improve/Rigor Explore para evidencia exploratoria acotada: validaciones en subconjunto, sweeps, búsqueda en GPU ociosa o transfer-learning rápido, con resúmenes sin sobreclamar en `explore_outputs/`. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)