Pptx Visual Assets
38.8kÚsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
- Costo de contexto al activarse
- 344 tok
- Tamaño del paquete
- 2 archivos
- Última actualización
- hace 26 días
Filtra checkpoints fine-tuneados con presupuestos de drift, comparación pareada y chequeos de olvido antes de promoverlos, tras un entrenamiento o al re-gatear un modelo ya promovido.
en todo el repo
0–100, la ruta de este skill
último commit aquí
últimos 90 días
66 tok en reposo
21 KB
Funciona con cualquier agente que lea SKILL.md
npx -y skills add wshobson/agents --skill checkpoint-promotion --agent claude-codeSe instala solo en este repositorio.
Di cualquiera de estas frases y el agente debería cargar este skill.
The Phase 5 gate for the whole
plugin: a checkpoint that trains
cleanly and beats its task metric
still doesn't ship without
clearing all four stages below.
eval-harness-first built the
suite re-run here — this skill is
where that suite's baseline
decides something.
Input: a trained checkpoint,
eval/baseline-<model>.json from
eval-harness-first, and the
frozen eval/drift-suite.yaml.
Output format:
promotion-report.md — the
four-stage evidence plus a
terminal PROMOTE or REJECT
verdict that /finetune Phase 5
and /promote-checkpoint consume
directly.
Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a deterministic arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.
trace-to-training-data's
Hygiene section exists to
prevent), and scan for label
noise. A checkpoint trained on
leaked goldens invalidates
every later stage.eval-harness-first's
eval/drift-suite.yaml —
MMLU/GSM8K/IFEval plus 200–500
domain-adjacent items — against
the checkpoint and diff against
baseline-<model>.json per
benchmark against the Drift
Budget table below.references/gate-templates.md
when every grader in the
harness is deterministic (no
LLM-judge; position
randomization N/A there).
A holdout win that
loses the live arena does not
ship — stage-2 numbers and
stage-3 judgments must agree; a
win on frozen goldens and a
loss in paired comparison is a
real signal, not a discrepancy
to explain away.| Drift (pts) | Verdict |
|---|---|
| ≤1 | Noise — proceed |
| 2–5 | Rerun with seed variation before deciding |
| >5 | HARD FAIL — no exception for task gains |
The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.
Item count derives from the
budget, not convenience: the
strict n for a half-width under
half the 5pt hard-fail threshold
is ~1,300 at typical accuracy
(p≈0.7); n=200 is a pragmatic
floor (±6pt half-width at that
same p, n=50 ±13pt) — report the
half-width with every verdict,
and treat a margin smaller than
it as REJECT (uncertain), not
PASS/HARD FAIL. Full math and a
5-run cautionary example:
references/gate-templates.md.
RERUN is not a verdict. A
2–5pt drift only ever produces a
PROMOTE or REJECT after the
seed-variation rerun completes —
PROMOTE requires landing back
at ≤1pt (noise); any rerun still
1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard
REJECT. No report may reach the Verdict section with stage 2 still showingRERUN.
Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:
If a checkpoint hits the >5pt
hard fail in stage 2, work this
escalation ladder in order — the
one canonical order this skill
and references/gate-templates.md
both point to:
lora-qlora-recipes and
preference-optimization tune
for the training run, applied
here in reverse.This order is a default, not a
law: remediation guidance from
a single before/after run pair
is a hypothesis — label it
low-confidence once any lever
produces a reversal, and prefer
a seed-variation repeat over
trusting the next rung blindly.
A lever that clears the drift
breach but drops a
success-criterion metric below
target is a two-sided tradeoff
for a human, not a reason to
keep descending the ladder. Full
reasoning and the 5-run
trajectory behind both caveats:
references/gate-templates.md.
Disclose drift-suite instruction reuse. A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.
promotion-report.md covers all
four stages as sections and
must end with a terminal
verdict: PROMOTE or REJECT,
the evidence that produced it,
and exactly one top remediation
when the verdict is REJECT.
Template: references/gate-templates.md.
The terminal contract other
skills parse:
## Verdict
REJECT
Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.
Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.
REJECT hands
the remediation back to a human
decision at
finetuning-method-selection or
the relevant training skill.eval-harness-first — owns the
drift suite and baseline this
skill re-runs and diffs
against; no baseline-<model>.json
means nothing to gate against.quantized-export — the only
valid next step after a
PROMOTE verdict.preference-optimization and
lora-qlora-recipes — own the
LR and rank levers in the
Catastrophic Forgetting
escalation path; this skill
diagnoses the breach, those
skills own the config that
caused it.dataset-curation — owns the
replay-mix construction recipe
the escalation ladder's first
rung applies.Complete promotion-report.md
template with all four stages,
the drift-suite scoring table,
the paired-arena protocol (item
count, position randomization,
win-rate threshold), and a
replay-mix configuration example:
references/gate-templates.md.
Reproducido de wshobson/agents bajo licencia MIT. Leer esta página en markdown.
2 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.
Requiere un checkpoint entrenado, eval/baseline-<model>.json de eval-harness-first y el eval/drift-suite.yaml congelado.
Este repo incluye 180 skills. Si instalas uno, normalmente ya tienes los demás.
Úsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción.
Úsalo al redactar o reparar una especificación JSON con coordenadas explícitas para un PPTX editable.
Úsalo para validar o reparar un PPTX editable en cuanto a geometría, accesibilidad, editabilidad nativa, linaje de fuente e integridad del paquete OOXML.
Úsalo para analizar un PPTX de referencia en modo solo lectura: estructura, tema, tipografía, ritmo de layout, diagnósticos, catálogos de plantillas derivados o inspección segura del paquete OOXML.
Úsalo al preparar la narrativa, las fuentes y el contexto de diseño para un nuevo deck PPTX editable.
Testea contratos inteligentes de forma exhaustiva con Hardhat y Foundry: tests unitarios, de integración y forking de mainnet.
Realiza auditorías de accesibilidad WCAG 2.2 con pruebas automatizadas, verificación manual y guía de remediación. Útil para auditar sitios, corregir violaciones y aplicar patrones de diseño accesible.
Prueba workflows de Temporal con pytest, time-skipping y estrategias de mocking: testing unitario, de integración, de replay y configuración de desarrollo local.