# Incident Runbook Templates > Crea runbooks estructurados de respuesta a incidentes con procedimientos paso a paso, rutas de escalamiento y acciones de recuperación para outages, bases de datos y onboarding de guardias. Fuente: https://skillsagentes.com/skills/wshobson/agents/incident-runbook-templates Markdown: https://skillsagentes.com/skills/wshobson/agents/incident-runbook-templates.md Repositorio: https://github.com/wshobson/agents Autor: wshobson Licencia: MIT Actualizado: hace 4 meses Coste de contexto: 121 tok instalada, 1.4k tok al activarse, 3.6k tok con todos los archivos del bundle Bundle: 2 archivos, 14 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add wshobson/agents --skill incident-runbook-templates --agent claude-code # Cursor npx -y skills add wshobson/agents --skill incident-runbook-templates --agent cursor # Codex npx -y skills add wshobson/agents --skill incident-runbook-templates --agent codex # Gemini CLI npx -y skills add wshobson/agents --skill incident-runbook-templates --agent gemini # Windsurf npx -y skills add wshobson/agents --skill incident-runbook-templates --agent windsurf # Cline npx -y skills add wshobson/agents --skill incident-runbook-templates --agent cline ``` ## Qué hace - Genera plantillas de runbooks de incidentes con detección, triaje, mitigación, resolución y comunicación - Define niveles de severidad (SEV1-SEV4) con tiempos de respuesta - Estructura matrices de escalamiento y checklists para responders bajo estrés - Incluye plantillas de comunicación con stakeholders durante incidentes ## Cuándo usarla - Crear procedimientos de respuesta a incidentes - Construir runbooks específicos de servicio - Establecer rutas de escalamiento y documentar recuperación - Incorporar ingenieros de guardia (on-call) ## Qué la activa - "Crea un runbook de outage para nuestro sistema de pagos" - "Necesito un procedimiento de incidente para agotamiento del connection pool" - "Ayúdame a estandarizar la matriz de escalamiento entre equipos" ## Antes de instalar - Necesita en el PATH: kubectl ## Archivos - SKILL.md — 5 KB - references/details.md — 9 KB ## SKILL.md Reproducido tal cual desde wshobson/agents bajo MIT. Esta sección es el documento original y está en inglés. # Incident Runbook Templates Production-ready templates for incident response runbooks covering detection, triage, mitigation, resolution, and communication. ## When to Use This Skill - Creating incident response procedures - Building service-specific runbooks - Establishing escalation paths - Documenting recovery procedures - Responding to active incidents - Onboarding on-call engineers ## Core Concepts ### 1. Incident Severity Levels | Severity | Impact | Response Time | Example | | -------- | -------------------------- | ----------------- | ----------------------- | | **SEV1** | Complete outage, data loss | 15 min | Production down | | **SEV2** | Major degradation | 30 min | Critical feature broken | | **SEV3** | Minor impact | 2 hours | Non-critical bug | | **SEV4** | Minimal impact | Next business day | Cosmetic issue | ### 2. Runbook Structure ``` 1. Overview & Impact 2. Detection & Alerts 3. Initial Triage 4. Mitigation Steps 5. Root Cause Investigation 6. Resolution Procedures 7. Verification & Rollback 8. Communication Templates 9. Escalation Matrix ``` ## Detailed patterns and worked examples Detailed pattern documentation lives in `references/details.md`. Read that file when the navigation tier above is insufficient. ## Best Practices ### Do's - **Keep runbooks updated** - Review after every incident - **Test runbooks regularly** - Game days, chaos engineering - **Include rollback steps** - Always have an escape hatch - **Document assumptions** - What must be true for steps to work - **Link to dashboards** - Quick access during stress ### Don'ts - **Don't assume knowledge** - Write for 3 AM brain - **Don't skip verification** - Confirm each step worked - **Don't forget communication** - Keep stakeholders informed - **Don't work alone** - Escalate early - **Don't skip postmortems** - Learn from every incident ## Troubleshooting ### Runbook steps work in staging but fail during a real incident Steps often assume preconditions that are true in a healthy environment but not during an outage. For each command in your runbook, add a prerequisite check and a "what to do if this command fails" note: ```bash # Step: Check pod status kubectl get pods -n payments # Prerequisites: kubectl configured, kubeconfig points to correct cluster # If this fails: run `aws eks update-kubeconfig --name prod-cluster --region us-east-1` # Expected output: pods in Running state ``` ### On-call engineer panics and skips steps out of order Add a numbered checklist at the top of the runbook that mirrors the section numbers, so responders can track progress under stress without reading the full document: ```markdown ## Quick Checklist - [ ] 1. Declare incident severity and open war room - [ ] 2. Check service health (Section 4.1) - [ ] 3. Check recent deployments (Section 4.1) - [ ] 4. Roll back if deploy is suspect (Section 4.1) - [ ] 5. Post initial notification to #payments-incidents - [ ] 6. Escalate if > 15 min unresolved ``` ### Runbook is outdated — commands reference old cluster names or endpoints Runbooks rot because they're updated manually. Include a "Last Verified" date and owner at the top, and add a CI check that validates all `curl` endpoints and `kubectl` context names are still valid: ```markdown ## Runbook Metadata | Field | Value | |---|---| | Last verified | 2024-11-15 | | Owner | @platform-team | | Review cadence | After every SEV1/SEV2 | ``` ### Stakeholder communication is delayed while engineers are heads-down Assign a dedicated incident communicator role (separate from the incident commander) whose only job is to post status updates. Add a standing agenda in the communication template: ``` Update every 15 minutes (even if no new information): - Current status (Investigating / Mitigating / Monitoring) - Impact (what is broken, who is affected, % of traffic) - What we are doing right now - Next update in: 15 minutes ``` ### Database runbook commands cause additional downtime when run incorrectly Add explicit warnings before destructive SQL commands and require a dry-run output check before executing: ```sql -- WARNING: This terminates active connections. Verify count first. -- DRY RUN (check count before terminating): SELECT count(*) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes'; -- EXECUTE only after verifying count is reasonable (< 50): SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes'; ``` ## Related Skills - `postmortem-writing` - After resolving an incident, use postmortem templates to capture root cause and preventive actions - `on-call-handoff-patterns` - Structure shift handoffs so the incoming responder has full context on active incidents ## Dónde encaja - Categoría: [DevOps e infraestructura](https://skillsagentes.com/categorias/devops-infraestructura.md) — Despliegues, contenedores, IaC y flujos de gestión de incidentes. - Creador: [wshobson](https://skillsagentes.com/creators/wshobson.md) — 183 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Hermes Tweet](https://skillsagentes.com/skills/wshobson/agents/hermes-tweet.md): Instala y opera Hermes Tweet, un plugin de Hermes Agent para investigar X/Twitter, leer timelines, analizar tweets y ejecutar operaciones privadas o de cambio de estado con aprobación previa. - [Superself](https://skillsagentes.com/skills/wshobson/agents/superself.md): Úsalo cuando un proyecto guarda su estado en Superself: lee `self context` al iniciar sesión, vincula el trabajo a una work unit, reporta con evidencia y registra decisiones confirmadas. - [Grounded Vault](https://skillsagentes.com/skills/wshobson/agents/grounded-vault.md): Úsalo para mantener un almacén Markdown de conocimiento donde cada afirmación compilada se rastrea hasta una fuente inmutable y el drift se detecta con git diff sin gastar tokens. - [Postgresql Table Design](https://skillsagentes.com/skills/wshobson/agents/postgresql-table-design.md): Úsalo al diseñar o revisar un esquema específico de PostgreSQL: buenas prácticas, tipos de datos, indexación, restricciones, patrones de rendimiento y funciones avanzadas. - [Prompt Engineering Patterns](https://skillsagentes.com/skills/wshobson/agents/prompt-engineering-patterns.md): Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción. ## Skills relacionadas - [Bazel Build Optimization](https://skillsagentes.com/skills/wshobson/agents/bazel-build-optimization.md): Optimiza builds de Bazel en monorepos a gran escala. Úsalo al configurar Bazel, implementar ejecución remota u optimizar el rendimiento de builds en codebases empresariales. - [Workflow Orchestration Patterns](https://skillsagentes.com/skills/wshobson/agents/workflow-orchestration-patterns.md): Diseña workflows duraderos con Temporal para sistemas distribuidos: separación workflow/activity, patrones saga, gestión de estado y restricciones de determinismo. - [Terraform Module Library](https://skillsagentes.com/skills/wshobson/agents/terraform-module-library.md): Construye módulos Terraform reutilizables para infraestructura AWS, Azure, GCP y OCI siguiendo buenas prácticas de infraestructura como código. - [Spark Training Gotchas](https://skillsagentes.com/skills/wshobson/agents/spark-training-gotchas.md): Preflight y diagnóstico de los diez fallos conocidos de entrenamiento ML en NVIDIA DGX Spark; úsala ante fallos de arranque, OOM bajo el límite de 128GB, caídas de throughput o antes de runs largos en GB10. - [Spark Environment Setup](https://skillsagentes.com/skills/wshobson/agents/spark-environment-setup.md): Configura un entorno de entrenamiento/inferencia de ML en NVIDIA DGX Spark (GB10, aarch64, CUDA 13): instalación de PyTorch/Unsloth/TRL/vLLM, errores ABI de libcudart y NGC vs pip. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)