Pptx Visual Assets
38.8kÚsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
- Costo de contexto al activarse
- 344 tok
- Tamaño del paquete
- 2 archivos
- Última actualización
- hace 26 días
Haz fine-tuning supervisado de modelos visión-lenguaje (VLMs) con datos de imagen+texto. Úsalo al adaptar un VLM a un dominio visual, configurar LoRA con torre de visión congelada, o depurar un fine-tune que entrena sin aprender.
en todo el repo
0–100, la ruta de este skill
último commit aquí
últimos 90 días
58 tok en reposo
14 KB
Funciona con cualquier agente que lea SKILL.md
npx -y skills add wshobson/agents --skill vision-sft --agent claude-codeSe instala solo en este repositorio.
Di cualquiera de estas frases y el agente debería cargar este skill.
This skill assumes finetuning-method-selection
already routed here: the data shape is
image+text demonstrations, not preference pairs
or a verifiable reward signal, and the base is a
vision-language model rather than a text-only
one. lora-qlora-recipes covers the text-only
LoRA/QLoRA recipe this skill specializes for the
vision tower and projector; read that skill first
if the LoRA fundamentals (rank, alpha, target
modules) aren't already familiar.
Input: an image+text dataset and a VLM base
model already picked from the model catalog.
Output format: a validated adapter config —
which components are frozen, LoRA target modules,
and a min_pixels/max_pixels budget — that
llm-finetuning-training-engineer consumes
directly when it generates a runnable script.
| Situation | Default |
|---|---|
| Adapting behavior on familiar images | Frozen tower+projector, LoRA r=8–16, α=16–32 |
| Visual domain shift | Unfreeze last-6 ViT layers, vision LR 5–10x lower |
| Doesn't fit in bf16 at target rank | QLoRA — frozen vision tower only |
fast_inference=True |
finetune_vision_layers=False |
| Loss normal, eval not improving | Check the Two Silent Killers below first |
Freeze the vision tower and the projector. Put
LoRA on the LLM only, all-linear (the same
attention + MLP target list as text-only SFT —
see lora-qlora-recipes), at r=8–16,
α=16–32. This is the settled default for
adapting a VLM's behavior without disturbing how
it sees.
# freeze tower + projector; LoRA on LLM only
for name, param in model.named_parameters():
if "vision_tower" in name or "projector" in name:
param.requires_grad = False
target_modules = [
"q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj",
] # LLM-only, all-linear — r=8-16, alpha=16-32
Unfreezing vision layers is a deliberate escalation, not a default decision — reach for it only when the domain shift is visual, not textual.
Both produce a run that trains without error and without learning: the loss curve looks normal, the model doesn't improve, and neither throws an exception — both need an explicit pre-training check, not just a clean training log.
references/collators-and-pitfalls.md.min_pixels/max_pixels resolution budget.
This pair is the single most consequential
hyperparameter for quality and memory in VLM
SFT — more than rank, alpha, or LR. Too low
silently downsamples images below what the task
needs (small document text becomes unreadable
even though training "succeeds"); too high blows
the activation memory budget or forces too small
a batch to train stably. Set it deliberately per
dataset, don't leave it at a framework default.UnslothVisionDataCollator is the collator
Unsloth expects for VLM SFT — it handles the
image-tag alignment and per-architecture
processor contract described in
references/collators-and-pitfalls.md. Don't
substitute a text-only collator for VLM data.finetune_vision_layers=False is required
when fast_inference=True. vLLM cannot serve
LoRA adapters on vision layers, so a fast-
inference setup that also unfreezes vision
layers fails at serve time even if training
succeeds. If the recipe calls for unfreezing the
last-6 ViT layers (see When to Unfreeze above),
fast inference is off the table for that run —
choose one or the other, not both.Base VLM choice is out of scope for this skill —
it lives in one place, the model catalog at
finetuning-method-selection's
references/model-catalog.md. This skill and its
references describe recipes by architecture
family only, never by recommending one model over
another.
VLM reinforcement learning (VLM-GRPO) is
reference-only in this plugin — the fragmented
tooling and reward-hacking failure modes specific
to VLM-RL are covered in grpo-rlvr-training,
not here. This skill's scope stops at supervised
fine-tuning.
The recurring mistake across every section above
is treating a clean loss curve as proof the run
is healthy. A normal-looking curve is consistent
with both a working run and either silent
killer, since the model trains on something
either way — just not the aligned image-text
signal when a killer is present. A flat eval score
next to a normal loss curve means re-run the
checklist in references/collators-and-pitfalls.md
before touching any hyperparameter.
references/collators-and-pitfalls.md — per-
architecture collator table, dataset-format
examples with image placeholders, a pre-
training validation checklist, and the two-
stage projector-alignment recipe as an advanced
pattern.Related skills: finetuning-method-selection
routes here; lora-qlora-recipes covers the
text-only LoRA fundamentals this skill
specializes; grpo-rlvr-training covers VLM-RL
(reference-only); dataset-curation covers
image+text dataset preparation this skill doesn't.
Reproducido de wshobson/agents bajo licencia MIT. Leer esta página en markdown.
2 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.
Asume que ya se eligió un modelo VLM base y que el dataset es de imagen+texto (no pares de preferencia ni recompensa verificable), tras pasar por `finetuning-method-selection`.
Este repo incluye 180 skills. Si instalas uno, normalmente ya tienes los demás.
Úsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.
Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción.
Úsalo al redactar o reparar una especificación JSON con coordenadas explícitas para un PPTX editable.
Úsalo para validar o reparar un PPTX editable en cuanto a geometría, accesibilidad, editabilidad nativa, linaje de fuente e integridad del paquete OOXML.
Úsalo para analizar un PPTX de referencia en modo solo lectura: estructura, tema, tipografía, ritmo de layout, diagnósticos, catálogos de plantillas derivados o inspección segura del paquete OOXML.
Úsalo al preparar la narrativa, las fuentes y el contexto de diseño para un nuevo deck PPTX editable.
Decide si conviene hacer fine-tuning y enruta al método correcto (SFT, DPO/ORPO/KTO, GRPO/RLVR, continued pretraining) y al modelo base adecuado.
Testea contratos inteligentes de forma exhaustiva con Hardhat y Foundry: tests unitarios, de integración y forking de mainnet.
Úsalo al seleccionar y colocar iconos, imágenes, SVGs, diagramas o infografías de apoyo aprobados en un PPTX editable.