Skills Agentes

Trace To Training Data

Convierte trazas de evaluación y logs de producción en ejemplos SFT y pares de preferencia; útil ante trazas calificadas, rejection sampling o construcción de pares DPO.

Estrellas
39.8k

en todo el repo

Actividad
52

0–100, la ruta de este skill

Actualizado
hace 2 meses

último commit aquí

Commits
1

últimos 90 días

Contexto
1.6k tok

69 tok en reposo

Paquete
2 archivos

14 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add wshobson/agents --skill trace-to-training-data --agent claude-code

Se instala solo en este repositorio.

Qué hace

  • Convierte trazas evaluadas y logs de producción en ejemplos SFT y pares de preferencia (DPO)
  • Selecciona la fracción de mayor reward entre las trayectorias exitosas para SFT
  • Construye pares DPO desde trayectorias pass/fail de la misma tarea usando selección μ−2σ
  • Aplica masking a nivel de paso en trayectorias multi-paso en lugar de descartarlas enteras
  • Escanea y redacta secretos/PII y evita que los goldens de eval se filtren al dataset de entrenamiento

Úsalo cuando

  • Existen trazas calificadas o ejemplos de fallo que deben convertirse en datos de entrenamiento
  • Se está aplicando rejection sampling a las salidas del modelo
  • Se están construyendo pares DPO a partir de runs exitosos y fallidos

No lo uses cuando

  • Una traza no tiene verdict ni reward todavía — hay que enviarla de vuelta a eval-harness-first en lugar de etiquetarla aquí

Qué lo activa

Di cualquiera de estas frases y el agente debería cargar este skill.

  • “Convierte estas trazas evaluadas de runs/results.json en ejemplos SFT”
  • “Genera pares DPO a partir de las trayectorias pass y fail de esta tarea”
  • “Aplica rejection sampling a estas salidas del modelo para curar el dataset”

SKILL.md

En inglés

Trace To Training Data

This skill assumes eval-harness-first already graded the traces being converted here — goldens, graders, and runs/<run-id>/results.json all exist before conversion starts. This is the flywheel edge that skill names in its own flow: "the same labeled traces become the training set." Conversion happens here; grading already happened upstream.

Input: graded traces — eval/goldens.jsonl plus runs/<run-id>/results.json, each row carrying a task_id, a verdict from the grader, and a reward when the task supports a scalar score (judge score, execution partial-credit, or an RLVR verifier):

{"task_id": "t-042", "trace_id": "t-042-a3",
 "messages": [{"role": "user", "content": "..."}],
 "verdict": "pass", "reward": 0.91,
 "grader": "exact_match"}

Output format: rows shaped exactly like dataset-curation's Format Selection table — SFT messages rows or DPO prompt/chosen/rejected pairs — so this skill's output is that skill's input with no reshaping step in between.

The Principle

The eval harness already did the labeling work: every trace in results.json carries a verdict, and often a reward, before this skill ever touches it. Converting a graded trace into a training row is mechanical — pick a shape from dataset-curation's table, map fields, write JSONL. Curation is the work that remains — which traces clear a quality bar, which pairs are informative, and which rows must never enter the training set at all.

Treat any conversion step that requires re-judging a trace as a sign the harness is missing a grader, not a gap this skill should paper over. A trace with no verdict or reward isn't convertible yet — route it back to eval-harness-first first, don't hand-label it here to unblock conversion.

SFT From Traces

  • Keep the top-reward fraction of successful trajectories, not every passing one. Rank passing traces by reward and take a fraction (the Agent-lightning pattern) rather than every trace that merely cleared the pass bar — a trace that barely passed is a weaker SFT signal than one that scored well above threshold.
  • Expert-corrected failures become gold SFT examples directly (the Langfuse pattern) — when a human edits a failing trace's output into a correct one, that correction needs no reward threshold; a human already validated it. Route corrections straight into the SFT set.
  • Step-level masking beats whole-trajectory discard for multi-step traces. When only some steps in a multi-step trajectory are bad, mask the loss on the bad steps and keep the good ones, rather than discarding the whole trajectory. SRFT reports 32.2% vs. 30.9% on SWE-bench for step-level critic masking over trajectory discard — a real, if modest, gap from the finer-grained cut.

Preference Pairs From Traces

  • Build pairs from passing-vs-failing trajectories on the SAME task, never from unrelated best- and worst-scoring traces pulled across different tasks — cross-task pairs teach the model to prefer one task over another, not one response over another.
  • Select the rejected member at μ−2σ of the reward distribution for that task, never the absolute minimum. preference-optimization's Pair Construction section owns the full selection formula; this skill supplies the graded trajectories it consumes.
  • Judge-scored delta selection cuts pair volume without cutting signal. Score each candidate pair by chosen-minus-rejected judge delta and keep only the highest-delta subset — the top 5k of a 16.5k candidate pool matched the full pool's downstream result. Build the full candidate set first, then filter by delta; don't cap generation at 5k up front.

Hygiene

  • Scan for secrets and PII before any row ships, and redact what's found. Traces sourced from production logs can carry credentials, API keys, tokens, or customer data — run a secret/PII scan over every SFT and DPO row and redact matches; conversion fails closed (the row is dropped, not shipped with the raw content) if sensitive fields remain after redaction. Never commit secrets.
  • Eval goldens must never leak into training data. Hold every eval/goldens.jsonl ID out of every converted SFT and DPO set — a trace that also appears as a golden trains on the exact item the checkpoint gets graded against later, silently inflating every subsequent eval run.
  • Dedup against the training set, not just within the newly converted rows — exact-match or embedding-similarity, matching dataset-curation's dedup method field, run against whatever training data already exists before this batch merges in.
  • Provenance goes into the dataset card. Every converted row must trace back to its source run_id and trace_id — dataset-curation's Provenance field checks for exactly this link back to trace-to-training-data output; a row with no traceable source isn't ready to merge.

Related Skills

  • eval-harness-first — produces the graded traces this skill converts; a trace with no verdict or reward isn't convertible yet, route it back there before conversion.
  • dataset-curation — owns the target formats and the dataset card this skill's provenance data feeds; converted rows must match its Format Selection table field names exactly, not an approximation of them.
  • preference-optimization — consumes the DPO pairs this skill builds and owns the full μ−2σ rejection-selection formula referenced above.

Worked JSONL-to-JSONL conversions — graded trace to SFT row, trace pair to DPO pair, correction to SFT row, the rejection-sampling loop, and the goldens-holdout check — live in references/conversion-recipes.md.

Reproducido de wshobson/agents bajo licencia MIT. Leer esta página en markdown.

Archivos

2 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

Antes de instalar

Requiere trazas ya calificadas por eval-harness-first, con eval/goldens.jsonl y runs/<run-id>/results.json existentes.

Detalles

Creador
wshobson
Licencia
MIT
Recursos incluidos
referencias
Repositorio
wshobson/agents
Código fuente
Ver SKILL.md

Etiquetas

Más de wshobson/agents

Este repo incluye 183 skills. Si instalas uno, normalmente ya tienes los demás. Ver el pack agents entero y su comando de instalación

Instala y opera Hermes Tweet, un plugin de Hermes Agent para investigar X/Twitter, leer timelines, analizar tweets y ejecutar operaciones privadas o de cambio de estado con aprobación previa.

Costo de contexto al activarse
1.4k tok
Tamaño del paquete
3 archivos
Última actualización
el mes pasado
redes sociales

Úsalo para mantener un almacén Markdown de conocimiento donde cada afirmación compilada se rastrea hasta una fuente inmutable y el drift se detecta con git diff sin gastar tokens.

Costo de contexto al activarse
1.4k tok
Tamaño del paquete
2 archivos
Última actualización
hace 24 días
documentos

Úsalo al diseñar o revisar un esquema específico de PostgreSQL: buenas prácticas, tipos de datos, indexación, restricciones, patrones de rendimiento y funciones avanzadas.

Costo de contexto al activarse
2k tok
Tamaño del paquete
2 archivos
Última actualización
hace 24 días
bases de datos

Úsalo cuando un proyecto guarda su estado en Superself: lee `self context` al iniciar sesión, vincula el trabajo a una work unit, reporta con evidencia y registra decisiones confirmadas.

Costo de contexto al activarse
1.3k tok
Tamaño del paquete
1 archivo
Última actualización
hace 24 días
productividad

Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción.

Costo de contexto al activarse
1.3k tok
Tamaño del paquete
10 archivos
Última actualización
el mes pasado
herramientas desarrollo

Audita y reescribe prosa para que deje de sonar generada por máquina. Incluye modo solo-detección, modo reescritura y modo edición en el lugar, con perfiles opcionales de voz y contexto.”

Costo de contexto al activarse
1.9k tok
Tamaño del paquete
4 archivos
Última actualización
el mes pasado
redaccion contenido

Skills relacionados

Construye sistemas de backtesting robustos para estrategias de trading, manejando correctamente look-ahead bias, survivorship bias y costes de transacción.

Costo de contexto al activarse
878 tok
Tamaño del paquete
2 archivos
Última actualización
hace 4 meses
datos analitica

Implementa validación de calidad de datos con Great Expectations, pruebas dbt y contratos de datos para pipelines fiables.

Costo de contexto al activarse
1.1k tok
Tamaño del paquete
2 archivos
Última actualización
hace 4 meses
datos analitica

Transforma datos en narrativas persuasivas usando visualización, contexto y estructura. Útil al presentar analíticas a stakeholders, crear reportes o presentaciones ejecutivas.

Costo de contexto al activarse
530 tok
Tamaño del paquete
2 archivos
Última actualización
hace 4 meses
datos analitica