# Avatar Video > Crea vídeos de avatar con IA controlando avatar, voz, guion, escenas y fondos mediante la API v2 de HeyGen, incluido WebM transparente e integración con Remotion. Fuente: https://skillsagentes.com/skills/calesthio/openmontage/avatar-video Markdown: https://skillsagentes.com/skills/calesthio/openmontage/avatar-video.md Repositorio: https://github.com/calesthio/OpenMontage Autor: calesthio Licencia: AGPL-3.0 Actualizado: hace 4 meses Coste de contexto: 140 tok instalada, 1.7k tok al activarse, 45.4k tok con todos los archivos del bundle Bundle: 16 archivos, 177 KB Permisos que pide: mcp__heygen__* ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add calesthio/OpenMontage --skill avatar-video --agent claude-code # Cursor npx -y skills add calesthio/OpenMontage --skill avatar-video --agent cursor # Codex npx -y skills add calesthio/OpenMontage --skill avatar-video --agent codex # Gemini CLI npx -y skills add calesthio/OpenMontage --skill avatar-video --agent gemini # Windsurf npx -y skills add calesthio/OpenMontage --skill avatar-video --agent windsurf # Cline npx -y skills add calesthio/OpenMontage --skill avatar-video --agent cline ``` ## Qué hace - Crea vídeos de avatar con la API v2 de HeyGen eligiendo avatar, voz y guion exactos - Monta vídeos multiescena con un fondo distinto por escena - Genera WebM transparente para composición posterior y admite fotos parlantes como presentador - Integra los avatares de HeyGen con Remotion y permite generación por lotes con especificaciones exactas ## Cuándo usarla - Quieres control preciso: eliges el avatar, escribes el guion exacto y configuras cada escena - Hace falta un vídeo multiescena con fondos distintos - Hace falta un WebM transparente para componer encima - Producción de vídeo con marca consistente y especificaciones exactas ## Cuándo no - El usuario solo describe una idea de vídeo y deja que la IA resuelva guion, avatar y visuales: para eso está el skill `create-video` ## Qué la activa - "Quiero que este avatar diga exactamente este guion" - "Hazme un vídeo multiescena con fondos distintos" - "Genera un WebM transparente con este presentador" - "Usa esta voz concreta para mi guion" ## Antes de instalar - Necesita `HEYGEN_API_KEY` y las herramientas `mcp__heygen__*`. - Necesita en el PATH: curl - Variables de entorno: HEYGEN_API_KEY - makes network requests - needs API credentials ## Archivos - SKILL.md — 6 KB - references/assets.md — 9 KB - references/avatars.md — 15 KB - references/backgrounds.md — 7 KB - references/captions.md — 6 KB - references/dimensions.md — 7 KB - references/photo-avatars.md — 24 KB - references/quota.md — 5 KB - references/remotion-integration.md — 18 KB - references/scripts.md — 10 KB - references/templates.md — 10 KB - references/text-overlays.md — 7 KB - references/video-generation.md — 22 KB - references/video-status.md — 13 KB - references/voices.md — 12 KB - references/webhooks.md — 9 KB ## SKILL.md Reproducido tal cual desde calesthio/OpenMontage bajo AGPL-3.0. Esta sección es el documento original y está en inglés. # Avatar Video Create AI avatar videos with full control over avatars, voices, scripts, scenes, and backgrounds. Build single or multi-scene videos with exact configuration using HeyGen's `/v2/video/generate` API. ## Authentication All requests require the `X-Api-Key` header. Set the `HEYGEN_API_KEY` environment variable. ```bash curl -X GET "https://api.heygen.com/v2/avatars" \ -H "X-Api-Key: $HEYGEN_API_KEY" ``` ## Tool Selection If HeyGen MCP tools are available (`mcp__heygen__*`), **prefer them** over direct HTTP API calls — they handle authentication and request formatting automatically. | Task | MCP Tool | Fallback (Direct API) | |------|----------|----------------------| | Check video status / get URL | `mcp__heygen__get_video` | `GET /v2/videos/{video_id}` | | List account videos | `mcp__heygen__list_videos` | `GET /v2/videos` | | Delete a video | `mcp__heygen__delete_video` | `DELETE /v2/videos/{video_id}` | Video generation (`POST /v2/video/generate`) and avatar/voice listing are done via direct API calls — see reference files below. ## Default Workflow 1. **List avatars** — `GET /v2/avatars` → pick an avatar, preview it, note `avatar_id` and `default_voice_id`. See [avatars.md](references/avatars.md) 2. **List voices** (if needed) — `GET /v2/voices` → pick a voice matching the avatar's gender/language. See [voices.md](references/voices.md) 3. **Write the script** — Structure scenes with one concept each. See [scripts.md](references/scripts.md) 4. **Generate the video** — `POST /v2/video/generate` with avatar, voice, script, and background per scene. See [video-generation.md](references/video-generation.md) 5. **Poll for completion** — `GET /v2/videos/{video_id}` until status is `completed`. See [video-status.md](references/video-status.md) ## Quick Reference | Task | Read | |------|------| | List and preview avatars | [avatars.md](references/avatars.md) | | List and select voices | [voices.md](references/voices.md) | | Write and structure scripts | [scripts.md](references/scripts.md) | | Generate video (single or multi-scene) | [video-generation.md](references/video-generation.md) | | Add custom backgrounds | [backgrounds.md](references/backgrounds.md) | | Add captions / subtitles | [captions.md](references/captions.md) | | Add text overlays | [text-overlays.md](references/text-overlays.md) | | Create transparent WebM video | [video-generation.md](references/video-generation.md) (WebM section) | | Use templates | [templates.md](references/templates.md) | | Create avatar from photo | [photo-avatars.md](references/photo-avatars.md) | | Check video status / download | [video-status.md](references/video-status.md) | | Upload assets (images, audio) | [assets.md](references/assets.md) | | Use with Remotion | [remotion-integration.md](references/remotion-integration.md) | | Set up webhooks | [webhooks.md](references/webhooks.md) | ## When to Use This Skill vs Create Video This skill is for **precise control** — you choose the avatar, write the exact script, configure each scene. If the user just wants to **describe a video idea** and let AI handle the rest (script, avatar, visuals), use the **create-video** skill instead. | User Says | Create Video Skill | This Skill | |-----------|:------------------:|:----------:| | "Make me a video about X" | ✓ | | | "Create a product demo" | ✓ | | | "I want avatar Y to say exactly Z" | | ✓ | | "Multi-scene video with different backgrounds" | | ✓ | | "Transparent WebM for compositing" | | ✓ | | "Use this specific voice for my script" | | ✓ | | "Batch generate videos with exact specs" | | ✓ | ## Reference Files ### Core Video Creation - [references/avatars.md](references/avatars.md) - Listing avatars, styles, avatar_id selection - [references/voices.md](references/voices.md) - Listing voices, locales, speed/pitch - [references/scripts.md](references/scripts.md) - Writing scripts, pauses, pacing - [references/video-generation.md](references/video-generation.md) - POST /v2/video/generate and multi-scene videos ### Video Customization - [references/backgrounds.md](references/backgrounds.md) - Solid colors, images, video backgrounds - [references/text-overlays.md](references/text-overlays.md) - Adding text with fonts and positioning - [references/captions.md](references/captions.md) - Auto-generated captions and subtitles ### Advanced Features - [references/templates.md](references/templates.md) - Template listing and variable replacement - [references/photo-avatars.md](references/photo-avatars.md) - Creating avatars from photos - [references/webhooks.md](references/webhooks.md) - Webhook endpoints and events ### Integration - [references/remotion-integration.md](references/remotion-integration.md) - Using HeyGen in Remotion compositions ### Foundation - [references/video-status.md](references/video-status.md) - Polling patterns and download URLs - [references/assets.md](references/assets.md) - Uploading images, videos, audio - [references/dimensions.md](references/dimensions.md) - Resolution and aspect ratios - [references/quota.md](references/quota.md) - Credit system and usage limits ## Best Practices 1. **Preview avatars before generating** — Download `preview_image_url` so the user can see the avatar before committing 2. **Use avatar's default voice** — Most avatars have a `default_voice_id` pre-matched for natural results 3. **Fallback: match gender manually** — If no default voice, ensure avatar and voice genders match 4. **Use test mode for development** — Set `test: true` to avoid consuming credits (output will be watermarked) 5. **Set generous timeouts** — Video generation often takes 5-15 minutes, sometimes longer 6. **Validate inputs** — Check avatar and voice IDs exist before generating ## Dónde encaja - Categoría: [Diseño y UI](https://skillsagentes.com/categorias/diseno-ui.md) — Sistemas de diseño, trabajo con componentes y acabado visual. - Creador: [calesthio](https://skillsagentes.com/creators/calesthio.md) — 0 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Seedance 2 5](https://skillsagentes.com/skills/calesthio/openmontage/seedance-2-5.md): Genera vídeo cinematográfico de 4-30 s con ByteDance Seedance 2.5 por fal.ai, Volcengine Ark, Runway o ComfyUI. Cubre el contrato de prompt 2.5, cortes duros, locks de continuidad y voz. - [Comfyui](https://skillsagentes.com/skills/calesthio/openmontage/comfyui.md): Úsalo al trabajar con workflows de ComfyUI en OpenMontage: comfyui_image/video/music, workflows propios, selección de output_node, modelos que faltan, LoRAs, poca VRAM e importación de workflows de la comunidad. - [Fish Audio Tts](https://skillsagentes.com/skills/calesthio/openmontage/fish-audio-tts.md): Genera narración expresiva y multilingüe con fish.audio (modelos S1 / S2) y reutiliza voces clonadas mediante reference_id. - [Minimax H3](https://skillsagentes.com/skills/calesthio/openmontage/minimax-h3.md): Genera vídeo con MiniMax H3 (Hailuo 3.0) por la API oficial v2, fal.ai, Runway, nodos partner de ComfyUI o pesos abiertos locales. Clips de 4-15s a 2K con animación de primer/último fotograma. - [Gemini Omni](https://skillsagentes.com/skills/calesthio/openmontage/gemini-omni.md): Genera y edita conversacionalmente vídeos cortos con Google Gemini Omni Flash: itera con ediciones en lenguaje natural, clips de 3-10s a 720p con audio y texto en pantalla, e imágenes de referencia por etiquetas. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)