# Saga Orchestration > Implementa patrones saga para transacciones distribuidas y workflows entre agregados, cuando 2PC no está disponible o hay que depurar sagas atascadas. Fuente: https://skillsagentes.com/skills/wshobson/agents/saga-orchestration Markdown: https://skillsagentes.com/skills/wshobson/agents/saga-orchestration.md Repositorio: https://github.com/wshobson/agents Autor: wshobson Licencia: MIT Actualizado: hace 4 meses Coste de contexto: 132 tok instalada, 1.5k tok al activarse, 7.8k tok con todos los archivos del bundle Bundle: 3 archivos, 31 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add wshobson/agents --skill saga-orchestration --agent claude-code # Cursor npx -y skills add wshobson/agents --skill saga-orchestration --agent cursor # Codex npx -y skills add wshobson/agents --skill saga-orchestration --agent codex # Gemini CLI npx -y skills add wshobson/agents --skill saga-orchestration --agent gemini # Windsurf npx -y skills add wshobson/agents --skill saga-orchestration --agent windsurf # Cline npx -y skills add wshobson/agents --skill saga-orchestration --agent cline ``` ## Qué hace - Diseña sagas de orquestación o coreografía con pasos ordenados y comandos de compensación - Genera lógica de compensación idempotente para cada servicio participante - Configura timeouts por paso y monitoreo de estados atascados - Aporta plantillas y patrones para recuperación desde DLQ ## Cuándo usarla - Coordinar transacciones multi-servicio sin locks distribuidos - Implementar transacciones compensatorias ante fallos parciales - Gestionar workflows de negocio de larga duración (minutos a horas) - Reemplazar el two-phase commit frágil por compensación asíncrona ## Qué la activa - "Necesito diseñar una saga para el flujo de reservas de hotel, vuelo y auto" - "Cómo implemento compensaciones para un pedido que falla en pago o envío" - "Mi saga se queda atascada en COMPENSATING, ayúdame a depurarla" ## Antes de instalar - Requiere infraestructura de mensajería existente (Kafka, RabbitMQ, SQS, etc.) y definir los límites de responsabilidad de cada servicio. ## Archivos - SKILL.md — 6 KB - references/advanced-patterns.md — 15 KB - references/details.md — 10 KB ## SKILL.md Reproducido tal cual desde wshobson/agents bajo MIT. Esta sección es el documento original y está en inglés. # Saga Orchestration Patterns for managing distributed transactions and long-running business processes without two-phase commit. ## Inputs and Outputs **What you provide:** - Service boundaries and ownership (which service owns which step) - Transaction requirements (which steps must be atomic, which can be eventual) - Failure modes for each step (transient vs. permanent, retry policy) - SLA requirements per step (informs timeout configuration) - Existing event/messaging infrastructure (Kafka, RabbitMQ, SQS, etc.) **What this skill produces:** - Saga definition with ordered steps, action commands, and compensation commands - Orchestrator or choreography implementation for your chosen pattern - Compensation logic for each participant service (idempotent, always-succeeds) - Step timeout configuration with per-step deadlines - Monitoring setup: state machine metrics, stuck saga detection, DLQ recovery --- ## When to Use This Skill - Coordinating multi-service transactions without distributed locks - Implementing compensating transactions for partial failures - Managing long-running business workflows (minutes to hours) - Handling failures in distributed systems where atomicity is required - Building order fulfillment, approval, or booking processes - Replacing fragile two-phase commit with async compensation --- ## Detailed section: Core Concepts Moved to `references/details.md`. ## Detailed section: Templates Moved to `references/details.md`. ## Best Practices ### Do's - **Make every step idempotent** — Commands may be replayed on broker reconnect - **Design compensations carefully** — They are the most critical code path - **Use correlation IDs** — The `saga_id` must flow through every event and log - **Implement per-step timeouts** — Never wait indefinitely for a participant reply - **Log state transitions** — `saga_id`, `step_name`, `old_state → new_state` on every change - **Test compensation paths explicitly** — Inject failures at each step index in integration tests ### Don'ts - **Don't assume instant completion** — Sagas are async and may take minutes - **Don't skip compensation testing** — The rollback path is the hardest to get right - **Don't couple services directly** — Use async messaging, never synchronous calls inside a saga step - **Don't ignore partial failures** — A step that partially executed still needs compensation - **Don't use a global timeout** — Each step has different latency characteristics --- ## Troubleshooting ### Saga stuck in COMPENSATING state A saga enters compensation but never reaches FAILED. This means a compensation handler is throwing an unhandled exception and never publishing `SagaCompensationCompleted`. Add dead-letter queue (DLQ) handling to compensation consumers and ensure every compensation action publishes a result event even when the underlying operation was already rolled back. ```python async def handle_release_reservation(self, command: Dict): try: await self.release_reservation(command["original_result"]["reservation_id"]) except ReservationNotFoundError: pass # Already released — treat as success # Always publish completion, regardless of outcome await self.event_publisher.publish("SagaCompensationCompleted", { "saga_id": command["saga_id"], "step_name": "reserve_inventory" }) ``` ### Duplicate saga executions on restart If your orchestrator service restarts mid-saga, it may replay events and re-execute already-completed steps. Guard every step action with an idempotency key — see **Template 3** above. ### Choreography saga losing events In a choreography-based saga, a downstream service may miss an event if it was offline when published. Use a durable message broker (Kafka with replication, RabbitMQ with persistence) and store the current saga state in a dedicated `saga_log` table so you can replay from the last known good step. ### Timeout firing before a slow-but-valid step completes A step like `create_shipment` might take up to 15 minutes during peak load but your global timeout is 5 minutes, causing spurious compensation. Make step timeouts configurable per step type — see `references/advanced-patterns.md` for the `TimeoutSagaOrchestrator` implementation and the `STEP_TIMEOUTS` dict pattern. ### Compensation order not matching execution order When two steps both complete before a failure is detected, compensation must run in strict reverse order or you leave data in an inconsistent state. Verify that `_compensate()` iterates from `current_step - 1` down to `0`, and add an integration test that deliberately fails at each step index to confirm correct rollback order. --- ## Advanced Patterns The `references/` directory contains production-grade implementations not needed for most sagas: - **`references/advanced-patterns.md`** — Full `SagaOrchestrator` abstract base class, `TimeoutSagaOrchestrator` with per-step deadlines, detailed bank transfer compensating transaction chain, Prometheus instrumentation, stuck saga PromQL alerts, and DLQ recovery worker. --- ## Related Skills - `cqrs-implementation` — Pair sagas with CQRS for read-model updates after each step completes - `event-store-design` — Store saga events in an event store for full audit trail and replay capability - `workflow-orchestration-patterns` — Higher-level workflow engines (Temporal, Conductor) that build on saga concepts ## Dónde encaja - Categoría: [DevOps e infraestructura](https://skillsagentes.com/categorias/devops-infraestructura.md) — Despliegues, contenedores, IaC y flujos de gestión de incidentes. - Creador: [wshobson](https://skillsagentes.com/creators/wshobson.md) — 183 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Hermes Tweet](https://skillsagentes.com/skills/wshobson/agents/hermes-tweet.md): Instala y opera Hermes Tweet, un plugin de Hermes Agent para investigar X/Twitter, leer timelines, analizar tweets y ejecutar operaciones privadas o de cambio de estado con aprobación previa. - [Superself](https://skillsagentes.com/skills/wshobson/agents/superself.md): Úsalo cuando un proyecto guarda su estado en Superself: lee `self context` al iniciar sesión, vincula el trabajo a una work unit, reporta con evidencia y registra decisiones confirmadas. - [Grounded Vault](https://skillsagentes.com/skills/wshobson/agents/grounded-vault.md): Úsalo para mantener un almacén Markdown de conocimiento donde cada afirmación compilada se rastrea hasta una fuente inmutable y el drift se detecta con git diff sin gastar tokens. - [Postgresql Table Design](https://skillsagentes.com/skills/wshobson/agents/postgresql-table-design.md): Úsalo al diseñar o revisar un esquema específico de PostgreSQL: buenas prácticas, tipos de datos, indexación, restricciones, patrones de rendimiento y funciones avanzadas. - [Prompt Engineering Patterns](https://skillsagentes.com/skills/wshobson/agents/prompt-engineering-patterns.md): Úsalo cuando pidan optimizar un prompt, mejorar su rendimiento, diseñar una plantilla, aplicar chain-of-thought, few-shot prompting o técnicas avanzadas de prompt engineering para producción. ## Skills relacionadas - [Bazel Build Optimization](https://skillsagentes.com/skills/wshobson/agents/bazel-build-optimization.md): Optimiza builds de Bazel en monorepos a gran escala. Úsalo al configurar Bazel, implementar ejecución remota u optimizar el rendimiento de builds en codebases empresariales. - [Workflow Orchestration Patterns](https://skillsagentes.com/skills/wshobson/agents/workflow-orchestration-patterns.md): Diseña workflows duraderos con Temporal para sistemas distribuidos: separación workflow/activity, patrones saga, gestión de estado y restricciones de determinismo. - [Terraform Module Library](https://skillsagentes.com/skills/wshobson/agents/terraform-module-library.md): Construye módulos Terraform reutilizables para infraestructura AWS, Azure, GCP y OCI siguiendo buenas prácticas de infraestructura como código. - [Spark Training Gotchas](https://skillsagentes.com/skills/wshobson/agents/spark-training-gotchas.md): Preflight y diagnóstico de los diez fallos conocidos de entrenamiento ML en NVIDIA DGX Spark; úsala ante fallos de arranque, OOM bajo el límite de 128GB, caídas de throughput o antes de runs largos en GB10. - [Spark Environment Setup](https://skillsagentes.com/skills/wshobson/agents/spark-environment-setup.md): Configura un entorno de entrenamiento/inferencia de ML en NVIDIA DGX Spark (GB10, aarch64, CUDA 13): instalación de PyTorch/Unsloth/TRL/vLLM, errores ABI de libcudart y NGC vs pip. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)