# Gemini Omni > Genera y edita conversacionalmente vídeos cortos con Google Gemini Omni Flash: itera con ediciones en lenguaje natural, clips de 3-10s a 720p con audio y texto en pantalla, e imágenes de referencia por etiquetas. Fuente: https://skillsagentes.com/skills/calesthio/openmontage/gemini-omni Markdown: https://skillsagentes.com/skills/calesthio/openmontage/gemini-omni.md Repositorio: https://github.com/calesthio/OpenMontage Autor: calesthio Licencia: AGPL-3.0 Actualizado: hace 9 días Coste de contexto: 154 tok instalada, 2.1k tok al activarse, 2.1k tok con todos los archivos del bundle Bundle: 1 archivo, 8 KB Permisos que pide: bash, read, write ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent claude-code # Cursor npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent cursor # Codex npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent codex # Gemini CLI npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent gemini # Windsurf npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent windsurf # Cline npx -y skills add calesthio/OpenMontage --skill gemini-omni --agent cline ``` ## Qué hace - Genera clips de 3 a 10 segundos a 720p con Gemini Omni Flash, con audio sintetizado y texto renderizado en pantalla - Permite editar el mismo clip por conversación en lenguaje natural, en capas, sin regenerarlo entero - Liga imágenes de referencia a roles con etiquetas `` e `` - Programa varios beats en un mismo prompt con sintaxis de timecode - Enruta la generación por `video_selector`, pero la edición se llama directa a `gemini_omni_video` porque el estado multi-turno vive ahí ## Cuándo usarla - Iterar sobre un clip con ediciones en lenguaje natural en vez de regenerarlo - Generar clips de 3-10s a 720p con audio o texto en pantalla - Ligar sujetos o estilos a imágenes de referencia con etiquetas en el prompt - Editar un vídeo ya subido ## Cuándo no - Clips cinematográficos de una sola pasada: mejor Seedance 2.0 - Clips de más de 10 segundos, por encima de 720p, o narración que no sea en inglés - Generaciones reproducibles por semilla, o interpolación de primer y último fotograma ## Qué la activa - "Haz el teléfono invisible y deja el resto igual" - "Genera un clip de 8 segundos con este texto en pantalla" - "Edita este vídeo subido y cámbiale el estilo" - "Genera un clip con estos beats por timecode" ## Antes de instalar - Necesita `GEMINI_API_KEY` o `GOOGLE_API_KEY`, la misma clave que usan Imagen y Google TTS; el modelo está en preview. - runs shell commands - writes to your files ## Archivos - SKILL.md — 8 KB ## SKILL.md Reproducido tal cual desde calesthio/OpenMontage bajo AGPL-3.0. Esta sección es el documento original y está en inglés. # Gemini Omni Flash (Google DeepMind) Gemini Omni is Google DeepMind's video generation **and editing** model family, announced at I/O 2026. The first model, **Gemini Omni Flash** (`gemini-omni-flash-preview`, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemini **Interactions API**. Its differentiator in the OpenMontage fleet is **stateful conversational editing**: each generation returns an `interaction_id`, and a follow-up call with `previous_interaction_id` edits that video in place — no other wrapped provider can refine a clip without regenerating it. OpenMontage wraps it as `gemini_omni_video` (native Gemini API, no gateway). It shares `GOOGLE_API_KEY`/`GEMINI_API_KEY` with `google_imagen` and `google_tts` — one key, three capabilities. Paid tier only: ~$0.10 per second of output video (billed as 5,792 output tokens/sec at $17.50/1M). Other documented routes are available when the direct Google key is not the chosen provider: | Route | OpenMontage call | Important limitation | |-------|------------------|----------------------| | fal.ai | `gemini_omni_fal` | T2V, I2V, reference video, and edit endpoints; no Google interaction ID is returned | | Runway | `runway_video`, `model: "gemini_omni_flash"` | T2V/I2V/V2V; video edits accept up to five image references | | ComfyUI Partner Node | `comfyui_video`, `model_family: "gemini_omni_flash"` | Hosted paid node; requires network, Comfy login, and credits | Use the direct `gemini_omni_video` route for stateful conversational editing. Gateway routes return ordinary provider tasks and cannot preserve Google's `previous_interaction_id` workflow. The fal edit endpoint can still be iterated by feeding each output video URL into the next edit call. ## When to pick it (and when not) | Use it for | Prefer another provider for | |---|---| | Iterative refinement — generate, review, then edit the same clip in layers | One-shot cinematic hero clips (→ Seedance 2.0, see `seedance-2-0`) | | Editing an existing/uploaded clip (restyle, add/remove objects, change text) | Clips longer than 10s or above 720p | | On-screen rendered text and word-by-word text beats | Seed-reproducible generations (no seed support) | | Reference-image-bound subjects/styles via prompt tags | First/last-frame interpolation (→ `veo_video`) | | Timecode-scheduled multi-beat clips from one prompt | Non-English narration (English only fully supported) | Route through `video_selector` for generation operations. **Editing (`edit_video`) is a direct-tool operation** — call `gemini_omni_video` from the registry, because the multi-turn interaction state lives outside the selector's model. ## Generation prompting Describe **scene + camera + lighting + motion + audio**. Official example: > Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air. - **Force a single shot** explicitly: "In a single continuous shot," / "No scene cuts." Otherwise the model may cut between scenes. - **Negatives go in prose** — there is no `negative_prompt` parameter: "No dialogue," "No extra sound effects." - **No sampler controls**: system instructions, temperature, top_p, and seeds are all unsupported. The prompt is the only lever. - **Meta-prompt for quality**: "Consider micro-detail, expression and timing to create a very rich, detailed but entirely natural scene." ### Timecode syntax Schedule beats with bracketed ranges or natural language — this maps directly onto OpenMontage scene-plan timings: ``` [0-3s] A person is walking [3-6s] They stop and turn around ``` > "After 3 seconds, a woman enters the scene." / "At 5s the chorus starts in the background audio." ### Audio and on-screen text Audio is synthesized automatically; direct it in the prompt: "Include calm background music," "The audio is a low tinny radio broadcast in the background." Rendered text works and can be timed: > One word on the screen at a time: 'did, you, know, that, Omni, can, do, awesome, text?' Each word appears for 1s. ## Reference images (`` / `` tags) Pass local images via `reference_image_paths` (they are sent in order), then bind them to roles **inside the prompt** with tags. `` indexes from 0 in the order supplied: ``` in the style of a woman is walking ``` ``` [0-3s] A studio fashion sequence. Starting with woman , she is holding [3-6s] Then we see the man holding ``` - `` makes an image the opening frame: ` a woman is walking`. - Use high-resolution images; describe the intended motion specifically rather than "make it move." - Say what each image *is* (product / character / style / background reference) — the model decides usage from context. ## Conversational editing (the differentiator) **Editing prompts are the opposite of generation prompts: short and surgical.** Overly descriptive edit prompts cause unintended changes. 1. Generate the base clip (subject + scene + motion). The tool returns `interaction_id` in its result data. 2. Pass it back as `previous_interaction_id` with `operation="edit_video"` and describe **only the delta**. 3. Append **"Keep everything else the same."** to pin unmentioned elements. 4. Refine in layers — one turn for lighting, one for camera, one for action, one for audio. Official good/bad pairs: | Avoid | Instead | |---|---| | "In the video of the man sitting on the sofa, please add a small black cat..." | "Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same." | | "Please remove the cell phone... and fill in the background so it looks like..." | "Make the phone invisible. Keep everything else the same." | Other working edit prompts: "Make this video anime" / "Put a fashionable hat on this person" / "Change the lighting to be more dramatic" / "Change the text on the sign to say 'Omni Flash'". **Gotcha — `store`:** editing via `previous_interaction_id` only works if the *prior* call kept the interaction server-side (`store` defaults to true in `gemini_omni_video`). Set `store=false` only for one-shot generations you will never edit. **Editing uploaded videos:** pass `input_video_path` instead of `previous_interaction_id`; the tool uploads it via the Files API. Unavailable in the EEA, Switzerland, and the UK (editing *generated* videos works everywhere). ## Hard limitations (preview) - Output: 3-10s, 720p, 24fps, MP4 with audio; aspect ratio `16:9` or `9:16`. All output carries an invisible SynthID watermark. - No seed, negative prompt, temperature, top_p, or system instructions. - No video extension or first/last-frame interpolation; no voice editing. - Audio reference inputs unsupported. Video references ≤3s are accepted by the schema but **not processed correctly** — don't rely on them. - Multi-video prompting unsupported; may degrade output. - English fully supported; other languages untested. - Images of minors (EEA/CH/UK) and certain recognizable people are blocked for upload/editing. ## Sources - Generation & editing guide: https://ai.google.dev/gemini-api/docs/omni - Model card: https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash - Pricing: https://ai.google.dev/gemini-api/docs/pricing - Announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/ ## Dónde encaja - Categoría: [Diseño y UI](https://skillsagentes.com/categorias/diseno-ui.md) — Sistemas de diseño, trabajo con componentes y acabado visual. - Creador: [calesthio](https://skillsagentes.com/creators/calesthio.md) — 0 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Seedance 2 5](https://skillsagentes.com/skills/calesthio/openmontage/seedance-2-5.md): Genera vídeo cinematográfico de 4-30 s con ByteDance Seedance 2.5 por fal.ai, Volcengine Ark, Runway o ComfyUI. Cubre el contrato de prompt 2.5, cortes duros, locks de continuidad y voz. - [Comfyui](https://skillsagentes.com/skills/calesthio/openmontage/comfyui.md): Úsalo al trabajar con workflows de ComfyUI en OpenMontage: comfyui_image/video/music, workflows propios, selección de output_node, modelos que faltan, LoRAs, poca VRAM e importación de workflows de la comunidad. - [Fish Audio Tts](https://skillsagentes.com/skills/calesthio/openmontage/fish-audio-tts.md): Genera narración expresiva y multilingüe con fish.audio (modelos S1 / S2) y reutiliza voces clonadas mediante reference_id. - [Minimax H3](https://skillsagentes.com/skills/calesthio/openmontage/minimax-h3.md): Genera vídeo con MiniMax H3 (Hailuo 3.0) por la API oficial v2, fal.ai, Runway, nodos partner de ComfyUI o pesos abiertos locales. Clips de 4-15s a 2K con animación de primer/último fotograma. - [Atlas Cloud](https://skillsagentes.com/skills/calesthio/openmontage/atlas-cloud.md): Genera o edita imágenes y vídeos por la pasarela Atlas Cloud: Seedance 2.5/2.0, Gemini Omni Flash, MiniMax H3, Seedream 5.0, GPT Image 2 y Nano Banana 2 con una sola ATLASCLOUD_API_KEY. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)