# Grpo Rlvr Training > Entrena comportamiento de razonamiento y tareas verificables con GRPO y RL a partir de recompensas verificables (RLVR): cuándo aplicar RL, la receta de referencia y el gate de inspección de rewards. Fuente: https://skillsagentes.com/skills/wshobson/agents/grpo-rlvr-training Markdown: https://skillsagentes.com/skills/wshobson/agents/grpo-rlvr-training.md Repositorio: https://github.com/wshobson/agents Autor: wshobson Licencia: MIT Actualizado: hace 2 meses Coste de contexto: 73 tok instalada, 2k tok al activarse, 5.8k tok con todos los archivos del bundle Bundle: 3 archivos, 23 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent claude-code # Cursor npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent cursor # Codex npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent codex # Gemini CLI npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent gemini # Windsurf npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent windsurf # Cline npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent cline ``` ## Qué hace - Configura y ejecuta entrenamiento GRPO con TRL usando rewards verificables (RLVR) - Define la receta base: GRPOConfig, num_generations≥8, learning_rate=5e-7, beta=0.01 - Impone la regla de inspección manual de 50-100 rewards antes del entrenamiento real - Guía la selección de variantes GRPO (DAPO, Dr.GRPO, GSPO) según el modo de fallo observado ## Cuándo usarla - El éxito de la tarea es algorítmicamente verificable (matemáticas, código, tool calls, salida estructurada) - Necesitas diseñar funciones de reward para GRPO - Un run de GRPO diverge o hace reward-hacking - El modelo ya tiene éxito parcial e inconsistente en la tarea objetivo ## Cuándo no - El modelo nunca tiene éxito ni a baja temperatura: es un problema de SFT, no de RL - La señal es una preferencia entre dos salidas aceptables (usa preference-optimization) - La evaluación requiere juicio humano o rúbrica subjetiva (usa eval-harness-first) ## Qué la activa - "Ayúdame a configurar un entrenamiento GRPO para verificar respuestas matemáticas" - "Mi run de GRPO está reward-hackeando, ¿cómo lo diagnostico?" - "Necesito diseñar una función de reward para validar salidas de tool calls" - "¿Qué variante de GRPO uso si el chain-of-thought colapsa en entropía?" ## Antes de instalar - Requiere una decisión de ruteo previa (RLVR vía GRPO) y un verificador (executor de código, test suite, validador de schema o grader). ## Archivos - SKILL.md — 8 KB - references/grpo-memory.md — 4 KB - references/reward-functions.md — 12 KB ## SKILL.md Reproducido tal cual desde wshobson/agents bajo MIT. Esta sección es el documento original y está en inglés. # GRPO & RLVR Training This skill assumes `finetuning-method-selection` already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (`lora-qlora-recipes`) or preference pairs (`preference-optimization`). What follows is when RL is the right tool, the reference recipe, the mandatory reward-inspection gate, and how to pick a GRPO variant when the base recipe misbehaves. **Input:** a routing decision (RLVR via GRPO) plus a verifier (code executor, test suite, schema checker, or grader) for the target task. **Output format:** a validated GRPO config — the kwarg values in `references/grpo-memory.md` and the reward functions in `references/reward-functions.md`, not free-form advice — that `llm-finetuning-training-engineer` consumes directly. ## When RL Applies GRPO+RLVR only pays off when task success is **algorithmically checkable** — a unit test passes, a parser accepts the output, a tool call matches an expected schema, a math answer matches a ground truth. If grading the output requires human judgment or a subjective rubric, that's an eval-harness and judge-calibration problem first — see `eval-harness-first` — not a reason to skip straight to RL. Before opening a GRPO run, confirm the model can **sometimes** succeed on the target task already. RL sharpens an existing capability by reweighting toward the samples that already work; it does not install a capability from zero. - **The model never succeeds, even at low temperature across many samples:** the gap is format or task understanding, not policy refinement. Route back to SFT first (`lora-qlora-recipes`) and only return to this skill once the base success rate is nonzero. - **The model succeeds sometimes, inconsistently:** this is the GRPO sweet spot — proceed to The Recipe below. The standing rule for the whole plugin: **DPO for taste, GRPO for reasoning.** If the signal is a preference between two acceptable outputs, that's `preference-optimization`, not this skill. ## The Recipe The reference recipe is TRL's `GRPOTrainer` with vLLM-backed generation: ```python from trl import GRPOConfig, GRPOTrainer grpo_args = GRPOConfig( output_dir="./outputs-grpo", use_vllm=True, vllm_mode="colocate", # single GPU; "server" for multi-GPU num_generations=8, # floor — fewer starves the group-relative baseline learning_rate=5e-7, # settled range for GRPO beta=0.01, # KL coefficient vs the reference policy per_device_train_batch_size=8, gradient_accumulation_steps=4, bf16=True, logging_steps=10, seed=3407, ) trainer = GRPOTrainer( model=SFT_CHECKPOINT, args=grpo_args, reward_funcs=[format_reward, correctness_reward], # references/reward-functions.md train_dataset=prompts, # prompt-only — GRPO generates its own completions processing_class=tokenizer, ) trainer.train() ``` - **`vllm_mode="colocate"`** runs generation and training on the same GPU — the default for a single-GPU box. - **`vllm_mode="server"`** points at a separate vLLM server process and is the multi-GPU path — generation and training don't compete for the same device. - **`num_generations` ≥ 8** is a floor, not a suggestion: GRPO's advantage estimate is relative to the group mean, and fewer than 8 samples per prompt produces a noisy baseline. - **Reward is composite** — a format reward (did the output parse / match the required structure) plus a correctness reward (did the answer verify). A well-formed-but-wrong answer and a malformed one should not score identically; correctness alone loses that signal. - **`learning_rate=5e-7`** and **`beta=0.01`** are the settled starting point; deviate only after the base run is stable and reward-inspected (below). Memory sizing for this recipe by target size class: `references/grpo-memory.md`. ## The Inspection Rule **Run the reward function against 50–100 sampled outputs and manually read the results before starting the actual training run.** This is a gate, not a one-time sanity check. If the reward function's judgment disagrees with a human reading of that sample, fix the reward function first. Training against an uninspected reward, or tuning hyperparameters to compensate for one silently scoring the wrong thing, is how a run reward-hacks: the model optimizes cleanly toward the wrong target, and that doesn't surface as a training-loop bug. This inspection is a Phase 1 gate input for `/finetune` — the same 50–100-sample read that catches a broken reward function here is what that command checks for before it lets a GRPO brief proceed. Complete reward function implementations to inspect against — exact-match, schema-validation, unit-test-execution, a length-penalty wrapper, and a rubric-as-reward judge pattern: `references/reward-functions.md`. ## Variant Selection The base recipe above is the default. Reach for a variant only when a specific failure mode shows up, not preemptively: | Failure mode | Variant | Why | |---|---|---| | Entropy collapse / degenerate long chain-of-thought | **DAPO** | Decouples clip bounds and relaxes the KL penalty that over-regularizes exploration on long reasoning traces | | Reward or output length trends up regardless of quality | **Dr.GRPO** | Removes GRPO's length-normalization bias so reward tracks correctness, not completion length | | Training a mixture-of-experts model | **GSPO** | Moves the importance-sampling ratio to the sequence level instead of per-token — per-token ratios are unstable on MoE routing, so GSPO is required here, not optional | Start with plain GRPO. Watch for the specific symptom — collapsing entropy on long CoT, a length-reward correlation, or MoE instability — and only then swap in the matching variant above. Don't pre-select a variant before the base recipe has actually shown the failure mode. ## VLM RL Is Reference-Only Vision-language RL is **not executed by this plugin in v1** — it's documented here for context, not as a runnable path. Tooling is fragmented across ms-swift and EasyR1-derived forks with no one-line TRL command yet, and naive text-only GRPO applied to a VLM tends to reward-hack by optimizing the text-reasoning trace while ignoring the image — the model learns to sound right without looking at the input. A VLM RL run is a research spike outside this skill's supported recipe, not a variant of The Recipe above. ## References - `references/reward-functions.md` — complete Python reward functions (exact-match correctness, schema validation, unit-test execution, a length-penalty wrapper, and a rubric-as-reward judge pattern) to inspect under The Inspection Rule before any training run. - `references/grpo-memory.md` — memory sizing by target size class, vLLM sleep-mode and optimizer-state tactics, Unsloth's long-context RL chunking, and the DGX Spark bandwidth caveat for decode-heavy rollouts. Related skills: `finetuning-method-selection` routes here once a verifiable pass/fail signal exists; `preference-optimization` is the sibling skill for preference pairs rather than verifiable rewards; `eval-harness-first` covers judge calibration for any reward that isn't purely code-checkable. On DGX Spark, defer to the `dgx-spark-ops` plugin's skills, when installed, for the memory/thermal remediation ladder this skill's memory table doesn't cover. ## Dónde encaja - Categoría: [Aprendizaje](https://skillsagentes.com/categorias/aprendizaje.md) — Explicaciones, tutorías y flujos de estudio estructurado. - Creador: [wshobson](https://skillsagentes.com/creators/wshobson.md) — 183 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Skills relacionadas - [Vision Sft](https://skillsagentes.com/skills/wshobson/agents/vision-sft.md): Haz fine-tuning supervisado de modelos visión-lenguaje (VLMs) con datos de imagen+texto. Úsalo al adaptar un VLM a un dominio visual, configurar LoRA con torre de visión congelada, o depurar un fine-tune que entrena sin aprender. - [Eval Harness First](https://skillsagentes.com/skills/wshobson/agents/eval-harness-first.md): Construye el eval harness que condiciona cada fine-tuning: golden sets, graders por modo de fallo, calibración de juez y baselines del modelo base. - [Review Agent Setup](https://skillsagentes.com/skills/wshobson/agents/review-agent-setup.md): Configura un gating humano para las acciones de revisión de agentes IA en Claude Code, con un rastro de aprobación auditable criptográficamente y gates aplicados con Cedar. - [Competitive Landscape](https://skillsagentes.com/skills/wshobson/agents/competitive-landscape.md): Analiza la competencia, identifica oportunidades de diferenciación y desarrolla posicionamiento de mercado ganador usando las Cinco Fuerzas de Porter, Blue Ocean Strategy y mapas de posicionamiento. - [Brand Landingpage](https://skillsagentes.com/skills/wshobson/agents/brand-landingpage.md): Diseñador de landing pages centrado en marca: hace una entrevista de identidad visual y luego genera e itera la página con Stitch, entregando HTML listo para desplegar. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)