Skills Agentes

Skill Autobench

Redacta un eval para una skill existente a partir de su historial real de uso (no de su spec), etiqueta casos como SPEC-DERIVED o HISTORY-IMPLIED y lo deja pendiente de aprobación humana.

Reemplaza a: Benchmarks derivados solo de la spec (--bootstrap-from-skill)

Estrellas
28.9k

en todo el repo

Actividad
59

0–100, la ruta de este skill

Actualizado
hace 9 días

último commit aquí

Commits
1

últimos 90 días

Contexto
3.1k tok

172 tok en reposo

Paquete
2 archivos

13 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add garrytan/gbrain --skill skill-autobench --agent claude-code

Se instala solo en este repositorio.

Qué hace

  • Mina invocaciones reales de una skill desde el archivo de conversaciones y transcripciones por harness, priorizando correcciones del usuario
  • Sintetiza un eval_contract y de 4 a 8 casos reproducibles con etiquetas HISTORY-IMPLIED o SPEC-DERIVED
  • Escribe el resultado en skills/<name>/eval/autobench-<date>.md con status: PENDING-HUMAN-APPROVAL, sin tocar SKILL.md
  • Verifica la integridad del panel multi-modelo (respuestas distintas, endpoints distintos, sin duplicados byte a byte)
  • Clasifica fallos minados según la taxonomía fail-improve (DETERMINISTIC-CODIFIABLE, PROMPT-FIXABLE, SPEC-GAP, ROUTING-MISS)

Úsalo cuando

  • Quieres escribir un eval para una skill existente basado en su uso real, no en su spec
  • Necesitas verificar que un panel de jueces multi-modelo realmente respondió con proveedores distintos
  • Quieres convertir correcciones de usuario repetidas en casos de prueba permanentes

No lo uses cuando

  • La skill objetivo no tiene ninguna historia de invocación real (usa gbrain skillopt <name> --bootstrap-from-skill en su lugar)

Qué lo activa

Di cualquiera de estas frases y el agente debería cargar este skill.

  • skill autobench para skill-x
  • escribe el eval a partir del historial de uso
  • sintetiza un eval para esta skill
  • verifica el panel de eval
  • ¿respondieron todos los proveedores?

SKILL.md

En inglés

skill-autobench — write the eval from lived usage

Convention: see conventions/brain-first.md — mining starts in the brain. Search the conversation archive before touching raw transcript files, and never declare "no history" without having queried the brain first.

Convention: see conventions/model-routing.md — mining and synthesis run on the cheap tier by default. The full multi-model judging pass is an explicit opt-in (see Contract).

The self-improving loop has three legs: an eval, a variant generator (SkillOpt), and a replay + judge harness (gbrain eval cross-modal). The generator and the judge ship with gbrain. The persistently missing leg is the eval author — someone has to WRITE the eval, and a spec-derived benchmark only tests what the skill promised, not what users actually asked for or what actually went wrong. This skill writes the eval from reality instead of imagination.

Pipeline

1. MINE — extract real invocation windows

Substrates, in priority order:

  1. Brain conversation archive — pages under conversations/, populated by the conversation-archive skill (hard dependency for this substrate: if it hasn't ingested your history yet, run it first). Search for the target skill's name, trigger phrases, and output shapes:

    gbrain search "<skill-name>"
    gbrain query "when did I use <skill-name> and what did I ask for"
    
  2. Per-harness session transcripts, as available — use what the harness exposes; do not assume a layout. gbrain transcripts recent --full reads the configured local transcript corpus (local-only by design). Claude Code keeps per-project session JSONL under ~/.claude/projects/; other harnesses have their own session stores. Absent stores are simply skipped.

From each hit, extract an invocation window: the user ask before the invocation, the invocation turn itself, and the 2 turns after — because that is where corrections live. A user correction after an invocation is the gold signal: it is a real, observed failure mode, and it becomes a hard_fail plus a replayable case.

FAIL-CLOSED: if no substrate yields a single real invocation of the target skill, emit an honest no-history report (substrates checked, queries run, windows scanned, zero matches) and stop. Do NOT invent "typical" invocations. For a skill with no history, the right tool is gbrain skillopt <name> --bootstrap-from-skill (spec-derived, and honest about it) — see Dedup.

2. SYNTH — turn windows into a proposed eval

From the mined windows plus the current SKILL.md, produce:

  • A proposed eval_contract: goal, dimensions, hard_fails. Dimensions come from observed asks; hard_fails encode observed corrections.
  • 4-8 replayable cases, each shaped {input, expected_behavior, failure_mode_to_catch} — realistic input, a checkable expected behavior, and the named failure mode the case exists to catch.
  • Spec-vs-usage gaps: "the spec says X, users consistently ask Y."

HONESTY LABELS are mandatory. Every dimension and every case is labeled:

  • HISTORY-IMPLIED — a real mined window backs it; cite which one.
  • SPEC-DERIVED — inferred from SKILL.md only; no usage evidence.

Never conflate the two. If history is thin or off-target, say so prominently at the top of the staged file ("GROUNDING WARNING: only N windows found, none exercised the core path") instead of padding with fabricated evidence.

Privacy scrub before staging: staged evals live in the skill repo and are distributable. Mined windows contain real names, companies, and deals — rewrite every case onto placeholder slugs (alice-example, acme-example) before writing the file. A history-grounded case keeps its shape and failure mode, never its real entities.

3. STAGE — human gate, always

Write skills/<name>/eval/autobench-<date>.md with frontmatter status: PENDING-HUMAN-APPROVAL.

This skill NEVER rewrites SKILL.md — not the eval_contract, not the body, not the triggers. Merging the staged eval is the human's decision. (This is a workflow contract the agent must honor, not a mechanically-enforced gate.)

The loop (after approval)

  1. Human reviews, edits, and approves the staged eval; the approved eval_contract is merged into the skill's frontmatter explicitly.

  2. Convert approved cases into skills/<name>/skillopt-benchmark.jsonl lines and run gbrain skillopt <name> — this is the SkillOpt surface extension: a history-grounded benchmark replacing the spec-derived bootstrap.

  3. Judge outputs through the native gate:

    gbrain eval cross-modal --task "<what the output was meant to achieve>" --output <path>
    
  4. Re-run autobench after more usage accumulates; diff against the prior staged baseline.

Panel integrity — trust no aggregate

Multi-model judging is only as good as the panel being real. The silent failure class: a model-id normalizer strips provider prefixes and every "different model" call lands on one host, so a "3-frontier consensus" is one model's opinion in a trench coat. Before trusting any multi-model verdict, assert over the result object:

  1. Each named model returned a non-empty response. An empty or failed slot is the first tell of a collapse.
  2. Responses came from DISTINCT provider endpoints. If three "different models" all report the same provider, the panel collapsed to one host.
  3. No two "different models" returned byte-identical output. If two differently-named models return the same bytes, they are the same model. This catches a collapse even when provider metadata is missing or faked.

gbrain eval cross-modal already exits 2 (INCONCLUSIVE) when fewer than 2/3 models return parseable scores; the byte-identical duplicate check and the distinct-endpoint check are the independent backstops this skill layers on top. Run them over the receipt JSON (written to the receipt dir) before treating a PASS/FAIL as authoritative. The integrity check is pure assertion logic over an existing result — it never calls a model itself: no network, no cost.

Fail-improve taxonomy — what mined failures become

Classify each mined correction/failure case by its cheapest durable fix:

Class Signal in history Durable fix
DETERMINISTIC-CODIFIABLE An LLM fallback repeatedly handles the same input shape (regex, parsing, slugs, dates) Convert to deterministic code + a permanent test case. The LLM is not the solution; it is the training-data generator for the code that replaces it.
PROMPT-FIXABLE The correction targets tone, format, or an omission the SKILL.md could specify An eval case + a gbrain skillopt run
SPEC-GAP Users consistently ask for something the spec never promised A spec-vs-usage gap observation for the human
ROUTING-MISS The skill fired on the wrong ask, or failed to fire A routing-eval.jsonl case, not a benchmark case

Direction of travel: every fixed failure becomes a permanent test, the deterministic share rises, and the LLM-fallback share falls. Log and improve; never silently drop a mined failure.

Contract

  • Input: a skill name that exists under skills/.
  • Substrates: conversations/ archive pages (brain-first), then per-harness session transcripts as available (gbrain transcripts recent is local-only by design). No usable substrate → honest no-history report, never fabricated evidence.
  • Output: one staged file at skills/<name>/eval/autobench-<date>.md, status: PENDING-HUMAN-APPROVAL — or the no-history report. Never an edit to SKILL.md, triggers, or any routing surface.
  • Cost posture: cheap-model default for mining and synthesis. The full multi-model judging pass (3 provider slots per cycle) is an explicit opt-in, and judging goes through native gbrain eval cross-modal — no bespoke judging harness.
  • Honesty: every dimension and case carries a SPEC-DERIVED or HISTORY-IMPLIED label; thin history is flagged, not papered over.
  • Privacy: mined cases are rewritten onto placeholder entities before staging.

Output Format

---
skill: <name>
status: PENDING-HUMAN-APPROVAL
generated: <date>
substrate: { conversation_pages: N, transcript_files: M, windows: K, corrections: C }
---

# Autobench: <name> — <date>

## Grounding
<one paragraph: how much real history backs this eval; GROUNDING WARNING if thin>

## Proposed eval_contract
goal / dimensions (each labeled HISTORY-IMPLIED|SPEC-DERIVED) / hard_fails

## Cases (4-8)
### case-01 [HISTORY-IMPLIED — window ref]
input: ...
expected_behavior: ...
failure_mode_to_catch: ...

## Spec-vs-usage gaps
- spec says X; users ask Y (windows: ...)

## Fail-improve classification
- case-03 → DETERMINISTIC-CODIFIABLE (same date-format fallback, 4 windows)

Anti-Patterns

  • ❌ Auto-merging a synthesized eval into SKILL.md. The human gate is the contract.
  • ❌ Presenting SPEC-DERIVED dimensions as history-grounded — fabricating usage evidence is the cardinal sin.
  • ❌ Mining nothing and still emitting a confident eval. Fail loudly or label honestly.
  • ❌ Trusting "3 models scored it 8/10" without checking that three providers actually returned distinct, non-identical responses.
  • ❌ Calling a model inside the panel-integrity check — it is pure assertion logic over a result object.
  • ❌ Building a bespoke judging harness when gbrain eval cross-modal is the native gate.
  • ❌ Staging mined cases with real people/companies in them. Placeholders only.

Dedup (sharp boundaries)

  • skill-optimizer (SkillOpt, host-side) — optimizes a skill's body against an EXISTING benchmark; its --bootstrap-from-skill derives tasks from the spec. THIS skill authors the benchmark from lived usage and feeds it into skills/<name>/skillopt-benchmark.jsonl — it extends SkillOpt's surface, never duplicates it. Read the skill-optimizer SKILL.md (on the host, where its engine lives) before the handoff. No history at all → use --bootstrap-from-skill, not this.
  • BrainBench (gbrain eval suites) — evals the ENGINE (retrieval, memory conformance, calibration). This evals SKILLS over their history.
  • skills/skillify/SKILL.md / skills/skill-creator/SKILL.md — create skills from descriptions; they don't mine lived usage.
  • skills/cross-modal-review/SKILL.md / gbrain eval cross-modal — RUN judging panels; they don't author evals. The panel-integrity assertions here verify their panels were real.
  • skills/skillpack-check/SKILL.md — audits skill structure/conformance, not behavior quality.
  • routing-eval.jsonl — tests dispatch (does the right skill fire); autobench tests behavior after dispatch. ROUTING-MISS findings route there.

Reproducido de garrytan/gbrain bajo licencia MIT. Leer esta página en markdown.

Archivos

2 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

Antes de instalar

Requiere el directorio conversations/ ya poblado (por la skill conversation-archive) para minar historial real.

Detalles

Creador
garrytan
Categoría
Testing y QA
Licencia
MIT
Recursos incluidos
Incluye scripts o referencias
Repositorio
garrytan/gbrain
Código fuente
Ver SKILL.md

Etiquetas

Más de garrytan/gbrain

Este repo incluye 75 skills. Si instalas uno, normalmente ya tienes los demás.

Setup

28.9k

Configura GBrain con auto-aprovisionamiento de Supabase o PGLite, inyección en AGENTS.md y primera importación.

Costo de contexto al activarse
7.4k tok
Tamaño del paquete
1 archivo
Última actualización
hace 4 días
bases de datos

Chequeos de salud del brain: aplicación de back-links, auditoría de citas, validación de filing, detección de info obsoleta, páginas huérfanas y benchmarks.

Costo de contexto al activarse
5k tok
Tamaño del paquete
1 archivo
Última actualización
hace 4 días
productividad

Migra un brain de gbrain-base a la taxonomía de 14 tipos canónicos de gbrain-base-v2 usando gbrain onboard --check y el handler Minion unify-types.

Costo de contexto al activarse
3.2k tok
Tamaño del paquete
1 archivo
Última actualización
hace 5 días
bases de datos

Cuándo y qué recuperar: abre la página del brain de una entidad relevante antes de responder desde memoria.

Costo de contexto al activarse
740 tok
Tamaño del paquete
1 archivo
Última actualización
hace 1 hora
productividad

Operaciones del brain: búsqueda primero, ciclo leer-enriquecer-escribir, atribución de fuentes, enriquecimiento ambiental y back-linking. Leer antes de cualquier interacción con el brain.

Costo de contexto al activarse
2.6k tok
Tamaño del paquete
1 archivo
Última actualización
hace 3 días
productividad

Importa exports de ChatGPT, Claude y Perplexity y transcripciones de sesiones como páginas fechadas en conversations/, valida y extrae hechos, y mantiene el archivo sin huecos con detección y backfill.

Costo de contexto al activarse
5k tok
Tamaño del paquete
2 archivos
Última actualización
hace 4 días
productividad

Skills relacionados

Control de calidad mediante un segundo modelo: hace que otro modelo revise el trabajo antes de darlo por bueno, con enrutamiento de negativas y opción de derivar a Codex para revisión de diffs.

Costo de contexto al activarse
1.7k tok
Tamaño del paquete
1 archivo
Última actualización
hace 3 meses
testing qa

Verificación sistemática, afirmación por afirmación, de cualquier contenido antes de publicarlo, basada en estándares de fact-checking profesional (The New Yorker, ProPublica, IFCN).

Costo de contexto al activarse
5.1k tok
Tamaño del paquete
2 archivos
Última actualización
hace 9 días
testing qa

Testing

28.9k

Framework de validación de skills más inteligencia diaria de salud y regresiones de la suite de tests: valida conformidad y clasifica fallos por tandas.

Costo de contexto al activarse
2k tok
Tamaño del paquete
1 archivo
Última actualización
el mes pasado
testing qa