ASD

Captions Overlay

Doctrina de overlay para captions embebidos: el modelo drop/rail/embed y la regla de que los captions son un overlay compuesto sobre el film, nunca una banda inferior reservada.

Reemplaza a: El keep-out gate mecánico de captions.mjs (retirado)

Estrellas
40.6k

en todo el repo

Actividad
57

0–100, la ruta de este skill

Actualizado
hace 28 días

último commit aquí

Commits
1

últimos 90 días

Contexto
1.5k tok

193 tok en reposo

Paquete
1 archivo

6 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add heygen-com/hyperframes --skill captions-overlay --agent claude-code

Se instala solo en este repositorio.

Qué hace

  • Define el modelo de captions: cada frase hablada es drop (filtrada), rail (subtítulo verbatim en primer plano) o embed (pico raro compuesto detrás del sujeto)
  • Prohíbe reservar una banda inferior vacía para captions; obliga a centrar la composición en el centro vertical real (y = H/2)
  • Establece que embed debe ser escaso: ≤1 por beat, nunca dos adyacentes o co-visibles, espaciados
  • Elimina el gate mecánico de keep-out; la legibilidad caption-sobre-contenido se juzga visualmente en QA

Úsalo cuando

  • Al añadir captions/subtítulos a un video de talking-head o de lanzamiento
  • Al decidir si una frase debe ser drop, rail o embed
  • Al maquetar una composición que llevará captions
  • Al centrar una composición en el centro real del frame bajo captions

No lo uses cuando

  • Para modo Cinematic con asks puramente cinemáticos donde no hace falta que las palabras se lean (se usa embed-style total, no rail)

Qué lo activa

Di cualquiera de estas frases y el agente debería cargar este skill.

  • Añade subtítulos a este video de talking-head
  • ¿Este video de lanzamiento reserva una banda para captions?
  • Decide qué frases deben ser rail y cuáles embed
  • Centra esta composición correctamente para que lleve captions

SKILL.md

En inglés

Captions Overlay Doctrine

Overlay doctrine — supplements the upstream embedded-captions skill. Applies ON TOP of it; do not expect it folded into the upstream skill.

Two ideas combine here. First, the caption model — every spoken phrase is drop, rail, or embed, and embed is the scarce earned peak, not the default. Second, the overlay law — a caption line is composited ON TOP of the film as an overlay; it is NOT a reserved zone, so you never shift content up or leave a dead band to "make room" for it. The two reinforce each other: because captions ride as an overlay (the verbatim rail in front, the occasional embed behind the subject), the composition keeps its full frame and centers on the true vertical center.

The caption model — drop / rail / embed

Every spoken phrase is one of three things (verbatim from embedded-captions):

What How it's shown
drop filler — um/uh, stutters, self-corrections not shown
rail the default — ordinary spoken content (verbatim) clean lower-third subtitle, in front, readable. A punch word can get an inline emphasis highlight (accent colour / active-word pop) — it stays on the rail.
embed a promoted peak — the headline beat one big word composited behind the subject (matte occlusion), designed entrance + exit

The rail carries most of the text; embed is the scarce, earned peak — ≤1 per beat, never two adjacent/co-visible, spaced ≥ a beat apart. A short clip → usually one embed; a long explainer → ~one per section. Embedding every word is the common mistake.

This is the Standard mode shape (rail = the verbatim lower-third; embed = the climax composited behind the subject). Cinematic mode drops the rail and makes everything embed-style — use it only for pure-cinematic asks, never for explainer / voiceover where the words must read.

Rail-first, embed-scarce (the load-bearing rules)

Quoted from the embedded-captions non-negotiables:

  • Rail-first for talking-head / explainer. Don't embed the whole transcript — most text is the rail; embed only peaks. Embedding everything is the default mistake.
  • Embed is scarce + spaced. ≤1 embed per sentence/beat, never two adjacent or co-visible, ≥ a beat apart, at most one apex. climax = per-beat peak, not "the single payoff of the entire clip."

The overlay law — captions are NOT a reserved band

In a generated launch composition, when captions are enabled, finalize composites a small, minimal word-by-word caption line as an overlay layer ON TOP of the whole film (a single text line, bottom-centered, roughly the bottom ~5-8% of canvas height). It is an overlay, not a reserved zone (verbatim from constraint #13 of the product-launch-video scene agent):

  • Center the composition on the TRUE vertical center — y = H / 2 (landscape 540, portrait 960). Do not shift content up to "make room" for captions; a composition centered at 0.42 × H with a dead lower band is the bug, not the fix.
  • Content may extend to the canvas bottom. Full-bleed subjects, rails, and backgrounds all welcome.
  • One soft courtesy rule: avoid parking critical small readable text (a URL line, a legal line, a sub-caption) exactly in the bottom ~80px center span where the caption line sits — the overlay would fight it. Large imagery / cards / ambient content under the captions is fine; the caption skin is designed to read over content.
  • There is no machine keep-out gate (the old captions.mjs keepout check is retired). Finalize snapshot QA judges caption-over-content legibility visually.

When captions are disabled: identical positioning freedom — the overlay simply doesn't exist.

Why these two rules are one doctrine

The model says the rail rides in front and an embed is a rare word composited behind the subject — both are layers added to footage that ships untouched. The overlay law says the caption line is a layer composited on top of the whole film, not a band carved out of the layout. So in both the captioning pipeline and the launch-video pipeline, captions are an overlay you add, not a zone you reserve:

  • Keep the full frame; center on true center; let content run to the edges.
  • Make the rail (or the small overlay caption line) carry the verbatim words.
  • Promote a word to an embed only at a genuine peak — scarce, spaced, never two at once.
  • Reserve nothing; judge legibility of captions-over-content visually, not by a keep-out gate.

Reproducido de heygen-com/hyperframes bajo licencia Apache-2.0. Leer esta página en markdown.

Archivos

1 archivo en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

Antes de instalar

Se aplica ENCIMA del skill embedded-captions, del cual cita el modelo rail+embed y la restricción #13 del scene agent de product-launch-video.

Detalles

Creador
heygen-com
Categoría
Diseño y UI
Licencia
Apache-2.0
Recursos incluidos
Solo SKILL.md
Código fuente
Ver SKILL.md

Etiquetas

Más de heygen-com/hyperframes

Este repo incluye 25 skills. Si instalas uno, normalmente ya tienes los demás.

Agent Media OS: resuelve BGM, SFX, imágenes, iconos, logos, voz, gradación de color o LUTs en un archivo local o bloque listo para usar; genera, opera y reutiliza assets en proyectos HyperFrames.

Costo de contexto al activarse
2k tok
Tamaño del paquete
152 archivos
Última actualización
hace 4 días
herramientas desarrollo

Convierte un pull request de GitHub (URL, owner/repo#N o 'este PR') en un video explicativo del cambio de código —changelog, feature, fix o refactor— construido a partir del diff, commits y archivos.

Costo de contexto al activarse
7.9k tok
Tamaño del paquete
30 archivos
Última actualización
hace 10 días
redaccion contenido

Punto de entrada obligatorio para crear, editar, animar o renderizar video, animaciones o motion graphics con HyperFrames, así como para inspeccionar, validar o publicar proyectos existentes.

Costo de contexto al activarse
3.2k tok
Tamaño del paquete
17 archivos
Última actualización
hace 15 días
diseno ui

Convierte una URL de producto o marketing, un guion pegado o un brief en un video de lanzamiento/promoción — promos SaaS, revelaciones de funciones, demos, lanzamientos de apps y empresas.

Costo de contexto al activarse
7.9k tok
Tamaño del paquete
28 archivos
Última actualización
hace 10 días
redaccion contenido

Usa el flujo de desarrollo de la CLI de HyperFrames: init, add, catalog, capture, lint, check, snapshot, compare, preview, render, publish, cloud, lambda, feedback, doctor y más; también para diagnosticar fallos de build o render.

Costo de contexto al activarse
3.9k tok
Tamaño del paquete
11 archivos
Última actualización
hace 3 días
herramientas desarrollo

Convierte texto (artículo, notas, tema o brief) en un video explicativo faceless: sin sitio ni material que capturar, los visuales se inventan por escena (tipografía, gráficos abstractos, diagramas, data-viz).

Costo de contexto al activarse
7.1k tok
Tamaño del paquete
24 archivos
Última actualización
hace 10 días
redaccion contenido

Skills relacionados

Empaqueta un video de talking-head/entrevista/podcast existente con tarjetas gráficas superpuestas (títulos, lower-thirds, callouts, citas, PiP) sincronizadas con la transcripción, en 16:9, 9:16 o 4:5; el clip se reproduce intacto debajo.

Costo de contexto al activarse
16.3k tok
Tamaño del paquete
28 archivos
Última actualización
hace 24 días
diseno ui

Porta el código de una composición Remotion (React) existente a HTML de HyperFrames. Solo para pedidos explícitos de portar/convertir/migrar una fuente Remotion, en un único sentido.

Costo de contexto al activarse
2.4k tok
Tamaño del paquete
70 archivos
Última actualización
hace 10 días
herramientas desarrollo

Úsalo cuando el usuario o el agente necesite leer, buscar o consultar la documentación o la referencia de la API de Stripe, en vez de usar curl o WebFetch para docs.stripe.com.

Costo de contexto al activarse
225 tok
Tamaño del paquete
1 archivo
Última actualización
hace 24 días
Oficialdesarrollo apis