Pptx Visual Assets
38.8kÚsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
- Costo de contexto al activarse
- 344 tok
- Tamaño del paquete
- 2 archivos
- Última actualización
- hace 26 días
Prepara, formatea y valida datasets para fine-tuning supervisado y entrenamiento de preferencias: formato, plantillas de chat, packing, datos sintéticos y dataset card.
en todo el repo
0–100, la ruta de este skill
último commit aquí
últimos 90 días
68 tok en reposo
27 KB
Funciona con cualquier agente que lea SKILL.md
npx -y skills add wshobson/agents --skill dataset-curation --agent claude-codeSe instala solo en este repositorio.
Di cualquiera de estas frases y el agente debería cargar este skill.
This skill assumes finetuning-method-selection
already routed here — the next step is preparing
data, not choosing a method. What follows: format
selection by target method, the template/packing
mechanics behind the most common silent training
failures, rules for mixing in synthetic data
without collapse, and the dataset card that closes
out Phase 2 before a run starts.
Input: raw examples (demonstrations, preference
judgments, or task prompts) plus a routing decision
from finetuning-method-selection.
Output format: a formatted, packed, validated
JSONL dataset plus a completed dataset card — the
Phase 2 artifact /finetune checks before launching
training.
| Method | Shape | Rows |
|---|---|---|
| SFT, single-turn | Instruct (instruction/response or prompt/completion) |
~1,000+ floor |
| SFT, multi-turn | Conversation / ChatML messages list |
~1,000+ floor |
| DPO / ORPO | Preference pair (prompt, chosen, rejected) |
Method-dependent, see preference-optimization |
| KTO | Unpaired (prompt, completion, label) |
Method-dependent, see preference-optimization |
| GRPO / RLVR | Prompt-only (prompt + verifier metadata) |
Method-dependent, see grpo-rlvr-training |
~1,000+ rows is the recommended floor for SFT, not a target. Below it, a handful of low-quality or duplicate examples can dominate the gradient; above it, quality over quantity — a smaller verified, deduplicated set beats a larger noisy one.
The ChatML shape, for orientation; the other four
formats plus a ShareGPT conversion note live in
references/formats-and-templates.md:
{"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]}
Apply the target model's chat template before any concatenation or packing, never after — packing raw text and templating the packed blob afterward corrupts turn boundaries, landing role markers in the wrong place relative to each example.
Train on assistant responses only. Mask the
loss (-100 in the labels tensor) over system/user
turns and the template's own role markers — only
assistant-turn content tokens contribute to loss.
Template/tokenizer mismatches are a top silent failure mode. A model trained against one chat template but served or evaluated with a different one degrades without erroring. Verify the same template string used in training is applied at inference and eval time.
Keep the dataset in messages shape and let
the trainer template and mask it
(assistant_only_loss=True in current TRL) —
pre-rendering to a flat text field destroys the
turn boundaries masking needs. Full code sketch:
references/formats-and-templates.md. Sanity-check
before training — decode only unmasked positions;
expect only assistant text:
keep = batch["labels"][0] != -100
print(tokenizer.decode(batch["input_ids"][0][keep]))
Without packing, 40–70% of compute is spent on padding — variable-length examples batched at a fixed sequence length waste the gap between each example's length and the batch's max. Packing concatenates multiple examples into one sequence up to the max length, cutting most of that waste.
Packing changes batch semantics. A packed sequence can contain several original examples, so "steps per epoch" and any LR schedule keyed to example count shift once packing is on — recompute schedule milestones against packed-sequence count.
MANDATORY: decode and manually inspect 5–10 packed sequences before scaling to a full run. Confirm example boundaries land where expected, template markers are intact per sub-example, and the loss mask is still assistant-only within each packed sequence. Not optional — packing bugs are silent (the loss curve looks normal) and only surface in eval quality, hours later:
for seq in packed_dataset.select(range(10)):
print(tokenizer.decode(seq["input_ids"]))
references/synthetic-data.md's
Replay-Mix Construction recipe);
state which rows count as "real"
in the dataset card rather than
leaving the floor structurally
unmeetable.references/synthetic-data.md.Every dataset that reaches training gets a card —
the required Phase 2 artifact /finetune checks
before launching. The card is not free-form
documentation; it MUST carry these fields:
trace-to-training-data output.references/synthetic-data.md.eval-harness-first run back to the checkpoint.A dataset missing any of these six fields isn't
ready for /finetune — the card is a gate, not a
summary written after the fact.
Before handing off to /finetune, confirm:
references/formats-and-templates.md — JSONL
examples per format, current-TRL masking code,
and the ShareGPT conversion note.references/synthetic-data.md — generation-method
ranking, filter funnel, replay-mix construction,
and teacher→student distillation pattern.Related skills: finetuning-method-selection routes
here; lora-qlora-recipes, vision-sft, and
preference-optimization consume the datasets this
skill produces; trace-to-training-data is the
provenance source for graded-trajectory datasets;
eval-harness-first grades the resulting checkpoint.
Reproducido de wshobson/agents bajo licencia MIT. Leer esta página en markdown.
3 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.
Asume que finetuning-method-selection ya se ejecutó y entregó una decisión de enrutamiento sobre el método a usar.
Este repo incluye 180 skills. Si instalas uno, normalmente ya tienes los demás.
Úsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción.
Úsalo al redactar o reparar una especificación JSON con coordenadas explícitas para un PPTX editable.
Úsalo para validar o reparar un PPTX editable en cuanto a geometría, accesibilidad, editabilidad nativa, linaje de fuente e integridad del paquete OOXML.
Úsalo para analizar un PPTX de referencia en modo solo lectura: estructura, tema, tipografía, ritmo de layout, diagnósticos, catálogos de plantillas derivados o inspección segura del paquete OOXML.
Úsalo al preparar la narrativa, las fuentes y el contexto de diseño para un nuevo deck PPTX editable.
Construye pipelines MLOps de extremo a extremo: preparación de datos, entrenamiento, validación y despliegue en producción del modelo.
Combina búsqueda vectorial y por palabras clave para mejorar la recuperación. Útil al implementar RAG, motores de búsqueda o cuando ningún enfoque solo basta.
Transforma datos en narrativas persuasivas usando visualización, contexto y estructura. Útil al presentar analíticas a stakeholders, crear reportes o presentaciones ejecutivas.