Skills Agentes

Video Generation

Se usa cuando el usuario pide generar, crear o imaginar videos. Admite prompts estructurados e imagen de referencia opcional para guiar la generación.

Estrellas
82.8k

en todo el repo

Actividad
48

0–100, la ruta de este skill

Actualizado
hace 4 meses

último commit aquí

Commits
1

últimos 90 días

Contexto
1.3k tok

37 tok en reposo

Paquete
2 archivos

13 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add bytedance/deer-flow --skill video-generation --agent claude-code

Se instala solo en este repositorio.

Este skill makes network requests, needs API credentials.

Qué hace

  • Genera videos de alta calidad a partir de prompts JSON estructurados y una imagen de referencia opcional, usando un script Python.
  • Puede usar la skill image-generation para crear primero una imagen de referencia que sirva de fotograma guía del video.
  • Selecciona automáticamente el proveedor (Gemini Veo o MiniMax) según las variables de entorno disponibles.

Úsalo cuando

  • El usuario pide generar, crear o imaginar videos.

No lo uses cuando

    Qué lo activa

    Di cualquiera de estas frases y el agente debería cargar este skill.

    • “Genera un video corto de la escena de apertura de esta película”
    • “Crea un clip de video a partir de esta imagen de referencia”
    • “Hazme un video de este producto con estilo cinematográfico”

    SKILL.md

    En inglés

    Video Generation Skill

    Overview

    This skill generates high-quality videos using structured prompts and a Python script. The workflow includes creating JSON-formatted prompts and executing video generation with optional reference image.

    Core Capabilities

    • Create structured JSON prompts for AIGC video generation
    • Support reference image as guidance or the first/last frame of the video
    • Generate videos through automated Python script execution

    Workflow

    Step 1: Understand Requirements

    When a user requests video generation, identify:

    • Subject/content: What should be in the image
    • Style preferences: Art style, mood, color palette
    • Technical specs: Aspect ratio, composition, lighting
    • Reference image: Any image to guide generation
    • You don't need to check the folder under /mnt/user-data

    Step 2: Create Structured Prompt

    Generate a structured JSON file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}.json

    Step 3: Create Reference Image (Optional when image-generation skill is available)

    Generate reference image for the video generation.

    • If only 1 image is provided, use it as the guided frame of the video

    Step 3: Execute Generation

    Call the Python script:

    python /mnt/skills/public/video-generation/scripts/generate.py \
      --prompt-file /mnt/user-data/workspace/prompt-file.json \
      --reference-images /path/to/ref1.jpg \
      --output-file /mnt/user-data/outputs/generated-video.mp4 \
      --aspect-ratio 16:9
    

    Parameters:

    • --prompt-file: Absolute path to JSON prompt file (required)
    • --reference-images: Absolute paths to reference image (optional)
    • --output-file: Absolute path to output image file (required)
    • --aspect-ratio: Aspect ratio of the generated image (optional, default: 16:9)

    [!NOTE] Do NOT read the python file, instead just call it with the parameters.

    Video Generation Example

    User request: "Generate a short video clip depicting the opening scene from "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe"

    Step 1: Search for the opening scene of "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe" online

    Step 2: Create a JSON prompt file with the following content:

    {
      "title": "The Chronicles of Narnia - Train Station Farewell",
      "background": {
        "description": "World War II evacuation scene at a crowded London train station. Steam and smoke fill the air as children are being sent to the countryside to escape the Blitz.",
        "era": "1940s wartime Britain",
        "location": "London railway station platform"
      },
      "characters": ["Mrs. Pevensie", "Lucy Pevensie"],
      "camera": {
        "type": "Close-up two-shot",
        "movement": "Static with subtle handheld movement",
        "angle": "Profile view, intimate framing",
        "focus": "Both faces in focus, background soft bokeh"
      },
      "dialogue": [
        {
          "character": "Mrs. Pevensie",
          "text": "You must be brave for me, darling. I'll come for you... I promise."
        },
        {
          "character": "Lucy Pevensie",
          "text": "I will be, mother. I promise."
        }
      ],
      "audio": [
        {
          "type": "Train whistle blows (signaling departure)",
          "volume": 1
        },
        {
          "type": "Strings swell emotionally, then fade",
          "volume": 0.5
        },
        {
          "type": "Ambient sound of the train station",
          "volume": 0.5
        }
      ]
    }
    

    Step 3: Use the image-generation skill to generate the reference image

    Load the image-generation skill and generate a single reference image narnia-farewell-scene-01.jpg according to the skill.

    Step 4: Use the generate.py script to generate the video

    python /mnt/skills/public/video-generation/scripts/generate.py \
      --prompt-file /mnt/user-data/workspace/narnia-farewell-scene.json \
      --reference-images /mnt/user-data/outputs/narnia-farewell-scene-01.jpg \
      --output-file /mnt/user-data/outputs/narnia-farewell-scene-01.mp4 \
      --aspect-ratio 16:9
    

    Do NOT read the python file, just call it with the parameters.

    Output Handling

    After generation:

    • Videos are typically saved in /mnt/user-data/outputs/
    • Share generated videos (come first) with user as well as generated image if applicable, using present_files tool
    • Provide brief description of the generation result
    • Offer to iterate if adjustments needed

    Notes

    • Always use English for prompts regardless of user's language
    • JSON format ensures structured, parsable prompts
    • Reference image enhance generation quality significantly
    • Iterative refinement is normal for optimal results

    Providers (Gemini / MiniMax)

    Auto-selected by environment variables (CLI unchanged):

    • GEMINI_API_KEY set → Gemini Veo (default, unchanged).
    • Only MINIMAX_API_KEY set → MiniMax video (/v1/video_generation, async 3-step poll/download).
    • Force with VIDEO_GENERATION_PROVIDER=gemini|minimax.

    MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com), MINIMAX_VIDEO_MODEL (default MiniMax-Hailuo-2.3). The first reference image is used as MiniMax first_frame_image. MiniMax ignores --aspect-ratio (it uses resolution/duration).

    Reproducido de bytedance/deer-flow bajo licencia MIT. Leer esta página en markdown.

    Archivos

    2 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

    Antes de instalar

    Requiere GEMINI_API_KEY o MINIMAX_API_KEY configurada; los prompts siempre se escriben en inglés.

    Necesita en el PATH:python

    Variables de entorno:GEMINI_API_KEYMINIMAX_API_HOSTMINIMAX_API_KEYMINIMAX_VIDEO_MODEL

    Detalles

    Creador
    bytedance
    Licencia
    MIT
    Recursos incluidos
    scripts en python
    Código fuente
    Ver SKILL.md

    Etiquetas

    Más de bytedance/deer-flow

    Este repo incluye 27 skills. Si instalas uno, normalmente ya tienes los demás. Ver el pack deer-flow entero y su comando de instalación

    Evalúa y ejecuta cambios de sistema no triviales desde primeros principios: RFCs, features, refactors, migraciones o nuevas APIs, exigiendo consumidores concretos, la solución mínima suficiente y evidencia proporcional al riesgo.

    Costo de contexto al activarse
    2.5k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 2 meses
    herramientas desarrollo

    Revisa paquetes de skills de DeerFlow: preparación para publicar, triggers, límites de seguridad, recursos y evidencia. Úsala cuando pidan auditar, calificar o validar para producción una skill.

    Costo de contexto al activarse
    1.4k tok
    Tamaño del paquete
    13 archivos
    Última actualización
    hace 2 meses
    herramientas desarrollo

    Crea skills nuevas, modifica y mejora skills existentes, y mide su rendimiento. Úsala para crear una skill, editarla, correr evals, hacer benchmark con análisis de varianza, u optimizar su descripción.

    Costo de contexto al activarse
    9.2k tok
    Tamaño del paquete
    20 archivos
    Última actualización
    hace 3 meses
    herramientas desarrollo

    Skill de smoke test de extremo a extremo para DeerFlow: actualiza el código, despliega en local o Docker, verifica disponibilidad de servicios, hace health check y genera el reporte final.

    Costo de contexto al activarse
    2.5k tok
    Tamaño del paquete
    12 archivos
    Última actualización
    hace 3 meses
    testing qa

    Manejo de issues y PRs de GitHub solo por comentarios para mantenedores de DeerFlow: resuelve alcance con gh, analiza, publica o redacta comentarios de issues y revisiones de PR, y compara PRs en competencia.

    Costo de contexto al activarse
    5.4k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    herramientas desarrollo

    Asegura que el código async del backend que podría bloquear el event loop de asyncio quede protegido por un test 'anchor' verificado en tests/blocking_io/, mediante un escaneo determinista de cambios o de todo el repo.

    Costo de contexto al activarse
    1.7k tok
    Tamaño del paquete
    4 archivos
    Última actualización
    hace 3 meses
    testing qa

    Skills relacionados

    Úsalo cuando el usuario pida revisar, analizar, criticar o resumir papers académicos, artículos de investigación, preprints o publicaciones científicas, con reviews estructuradas al estilo de las mejores conferencias.

    Costo de contexto al activarse
    3k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 6 meses
    investigacion

    Asegura que el código async del backend que podría bloquear el event loop de asyncio quede protegido por un test 'anchor' verificado en tests/blocking_io/, mediante un escaneo determinista de cambios o de todo el repo.

    Costo de contexto al activarse
    1.7k tok
    Tamaño del paquete
    4 archivos
    Última actualización
    hace 3 meses
    testing qa

    Genera un SOUL.md personalizado mediante una conversación de onboarding cálida y adaptativa. Se activa al crear, configurar o inicializar la identidad de un compañero de IA.

    Costo de contexto al activarse
    1.2k tok
    Tamaño del paquete
    3 archivos
    Última actualización
    hace 5 meses
    productividad