# Convex Self Heal > Convierte un error de producción en un PR de fix triado, con causa raíz identificada, reparado y certificado (tsc + rehearsal + reproduce-then-gone) para que un humano lo mergee — nunca hace merge automático. Fuente: https://skillsagentes.com/skills/get-convex/agent-skills/convex-self-heal Markdown: https://skillsagentes.com/skills/get-convex/agent-skills/convex-self-heal.md Repositorio: https://github.com/get-convex/agent-skills Autor: get-convex Licencia: Apache-2.0 Actualizado: el mes pasado Coste de contexto: 49 tok instalada, 1.4k tok al activarse, 1.4k tok con todos los archivos del bundle Bundle: 1 archivo, 5 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent claude-code # Cursor npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent cursor # Codex npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent codex # Gemini CLI npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent gemini # Windsurf npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent windsurf # Cline npx -y skills add get-convex/agent-skills --skill convex-self-heal --agent cline ``` ## Qué hace - Toma un error de producción y lo triaja, root-causea, repara en una rama y certifica el fix (tsc + rehearsal + reproduce-then-gone) antes de abrir un PR - Orquesta sentinel (captura) → findings bus (diagnóstico) → fixers (repara) → migrate-rehearse/tsc/probe (certifica) → PR humano (decide) → deploy-guard (promueve) - Clasifica errores como transitorios, de config o defectos de código/schema, y solo repara los últimos - Tras el merge y deploy, revisa la tabla sentinel + logs para confirmar que el error deja de repetirse, y reabre si persiste - Solo auto-prepara clases de fix pre-aprobadas (validator, índice, ownership-check, backfill no destructivo); deriva lo destructivo o ambiguo a un humano ## Cuándo usarla - Aparece un error nuevo o recurrente en producción capturado por sentinel y hace falta triaje, root-cause y fix certificado - Se necesita un PR de reparación con evidencia de certificación (tsc, rehearsal, reproduce-then-gone) para que un humano lo revise y mergee - Después de un merge, hay que confirmar que la firma de error dejó de aparecer en prod ## Cuándo no - No hay sentinel instalado (el skill ofrece instalarlo y se detiene, no hay nada que curar sin captura) - El error es un blip transitorio de red (se reintenta/ignora, no se abre PR) - La causa raíz es incierta (se detiene y reporta en vez de proponer un fix no certificado) ## Qué la activa - "Tenemos un error recurrente en sentinel, investígalo y prepara un fix certificado" - "Revisa este error de producción y abre un PR con la reparación si se puede certificar" - "Confirma si el error de la última semana dejó de aparecer tras el deploy" ## Antes de instalar - Requiere sentinel instalado para capturar errores de producción en el propio deployment del usuario. ## Archivos - SKILL.md — 5 KB ## SKILL.md Reproducido tal cual desde get-convex/agent-skills bajo Apache-2.0. Esta sección es el documento original y está en inglés. # Gated production self-healing loop Sentry/Datadog/Vercel can go error→investigate→draft-PR, but they treat the backend as opaque and stop at the human merge gate with an unverified diff. Convex can do the step they can't: because the error rows live in the user's own deployment and the fix can be rehearsed on a preview of that deployment, the platform certifies the fix against real invariants before anyone reviews it. This capability is the composition capstone — it wires sentinel (capture) → the findings bus (diagnose) → the fixers (repair) → migrate-rehearse/tsc/probe (certify) → a human PR (decide) → deploy-guard (promote). The human keeps the merge button; the machine does everything up to and including proving the fix works. ## Workflow 1. GUARD: deploy-guard — this loop reads prod and PROPOSES prod changes; classify + announce the deployment and get the standing consent for the loop's scope up front (what classes of fix it may auto-prepare vs must always defer). Never auto-merge; the human merge is the fixed boundary. 2. CAPTURE: require sentinel (prod errors in the user's own deployment, redacted at write time). If absent, offer to install it and stop — there is nothing to heal without capture. 3. TRIAGE a new/ recurring error: pull it via the official MCP (data/run-once-query over the sentinel table, or the monitor's prod_error event). Classify: transient (retry/ignore — do NOT open a PR for a one-off network blip), config (env/secret — hand to env, never guess a secret), or a code/schema defect (proceed). 4. ROOT-CAUSE on the findings bus: run the relevant audit pass on the implicated function — convex-insights (the failing requests + stacks), convex-advisor (if it's a read-limit/OCC cause), convex-reviewer/convex-authz (if it's a logic/authz defect). Produce a bus finding with evidence (the stack + the reproducing input) and a fixCapability. If root cause is unclear, STOP and report — a wrong fix is worse than an open error. 5. REPAIR via the finding's fixCapability (convex-authz, reviewer fixers, convex-expert for perf) on a branch — never on prod directly. 6. CERTIFY against the backend's own invariants BEFORE proposing (this is the differentiator — do not skip any that apply): (a) `tsc --noEmit` clean; (b) if the fix touches schema/data, run it through migrate-rehearse on a preview seeded with a prod snapshot — the schema-conformance gate must pass on real-shaped data; (c) reproduce-then-confirm-gone: replay the error's triggering input against the fixed code (a convex-test case or an MCP run on the preview) and assert the failure no longer occurs; (d) no-regression: the finding must be gone AND no new bus finding introduced on the touched function. A fix that fails any applicable certification is NOT proposed — it's reported as 'attempted, could not certify' with what failed. 7. PROPOSE, never merge: open a PR (or a diff for review) containing the fix, the certification evidence (tsc result, rehearsal outcome, the reproduced-then-gone assertion), the original error + finding, and the reversibility note. Label the change class. The human reviews and merges. 8. PROMOTE on merge via deploy-guard's prod consent; after deploy, re-check the sentinel table + `logs` (failures) to confirm that error signature stops recurring (do NOT use `insights` for this — it tracks only OCC/read-limit perf events, not arbitrary error signatures) — the loop is only closed when the error stops recurring in prod. If it recurs, reopen with the new evidence. 9. BOUND it: only classes the user pre-approved in step 1 are auto-prepared (default-safe set: validator fixes, missing-index adds, ownership-check adds, non-destructive backfills); anything destructive, security-sensitive beyond an added check, or ambiguous is always deferred to explicit human direction. Log every action to an append-only record so the loop is auditable. ## Rules - The human keeps the merge button — this loop prepares and certifies fixes, it NEVER auto-merges or auto-deploys to prod (matches the industry boundary: no credible system ships unattended prod auto-merge). - Certify before proposing: tsc + (schema→migrate-rehearse on a prod-snapshot preview) + reproduce-then-confirm-the-failure-is-gone + no new bus finding. An uncertified fix is reported as 'could not certify', never proposed as done. - Triage first: transient blips get retried/ignored, config errors go to env (never guess a secret), only real code/schema defects enter the repair loop. - Repair on a branch/preview, never on prod directly; promote only through deploy-guard's fresh prod consent. - Only pre-approved fix classes are auto-prepared (default-safe: validator/index/ownership/non-destructive backfill); destructive or ambiguous changes are always deferred to the human. - Close the loop for real: after merge+deploy, confirm the error signature stops recurring via the sentinel table + logs (not insights, which only sees perf events); reopen if it persists. - Every action is logged to an append-only, auditable record; data residency stays in the user's own deployment (sentinel discipline). - If root cause is unclear, STOP and report — an uncertain fix is worse than an open, visible error. ## Dónde encaja - Categoría: [DevOps e infraestructura](https://skillsagentes.com/categorias/devops-infraestructura.md) — Despliegues, contenedores, IaC y flujos de gestión de incidentes. - Creador: [get-convex](https://skillsagentes.com/creators/get-convex.md) — 33 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)