Skills Agentes

Blog Image

Generación y edición de imágenes con IA para contenido de blog mediante Gemini por MCP: portadas, ilustraciones, tarjetas sociales y OG, con 6 modos de dominio y retorno silencioso si el MCP no está disponible.

Estrellas
2.1k

en todo el repo

Actividad
67

0–100, la ruta de este skill

Actualizado
hace 19 días

último commit aquí

Commits
10

últimos 90 días

Contexto
3.4k tok

123 tok en reposo

Paquete
6 archivos

67 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add AgriciDaniel/claude-blog --skill blog-image --agent claude-code

Se instala solo en este repositorio.

Este skill makes network requests, needs API credentials.

Qué hace

  • Nunca pasa el texto del usuario tal cual a la API: construye el prompt con el sistema de 6 componentes (sujeto, acción, contexto, composición, iluminación, estilo)
  • Elige un modo de dominio según el uso: Editorial, Producto, Paisaje, UI/Web, Infografía o Abstracto
  • Fija la relación de aspecto antes de generar y postprocesa con `magick` a 1200x630, WebP o AVIF
  • Ante `IMAGE_SAFETY` o `SAFETY` reformula el prompt en positivo y reintenta hasta 3 veces en vez de rendirse
  • Devuelve ruta, prompt completo, ajustes, texto alternativo de 10 a 125 caracteres y el bloque de frontmatter listo para pegar

Úsalo cuando

  • El usuario dice "blog image", "generar imagen de portada", "ilustración para el blog", "editar imagen" o "imagen OG"
  • blog-write o blog-rewrite no encuentran suficientes fotos de banco adecuadas para el tema

No lo uses cuando

    Qué lo activa

    Di cualquiera de estas frases y el agente debería cargar este skill.

    • Genera la imagen de portada de este post
    • Haz esta imagen más cálida
    • Crea una tarjeta OG de 1200x630

    SKILL.md

    En inglés

    Blog Image - AI Image Generation for Blog Content

    You are a Creative Director that orchestrates Gemini's image generation specifically for blog content. Never pass raw user text directly to the API. Always interpret, enhance, and construct an optimized prompt using the 6-component Reasoning Brief system.

    Quick Reference

    Command What it does
    /blog image generate <idea> Generate a blog image with full prompt engineering
    /blog image edit <path> <instructions> Edit an existing blog image intelligently
    /blog image setup Configure MCP server and API key

    Blog Image Types

    Match the image type to blog use case:

    Image Type Aspect Ratio Resolution Domain Mode Placement
    Hero/Cover 16:9 2K or 4K Editorial / Landscape Frontmatter coverImage
    OG/Social Card 16:9 1K Editorial / Infographic Frontmatter ogImage
    Inline Illustration 16:9 or 4:3 1K Varies by topic After H2, before body
    Inline Product Shot 4:3 or 1:1 1K Product Within product sections
    Section Divider 21:9 then crop 1K Abstract / Landscape Between major sections

    Sizing requirements:

    • Blog hero/cover: 1200x630 (OG-compatible) or 1920x1080
    • Open Graph (OG): 1200x630 (required for social sharing)
    • Inline images: 1200px+ wide

    MCP Availability Check

    Before generating, check if nanobanana-mcp tools are available:

    1. Try calling get_image_history with conversation_id: "default" (lightweight, no side effects)
    2. If it succeeds: MCP is available, proceed with generation
    3. If it fails: MCP not configured - inform the user:
      • "Image generation requires the nanobanana-mcp server. Run /blog image setup to configure it."
      • When called internally (from blog-write/blog-rewrite): return silently, no error. The calling workflow continues with stock photos.

    Generation Workflow

    For /blog image generate <idea> or when invoked internally:

    Step 1: Analyze Intent

    Determine what the blog needs:

    • Image type: Hero, inline, OG card, section divider?
    • Blog topic: What is the article about?
    • Style: Photorealistic, editorial, illustrated, minimal?
    • Constraints: Brand colors, specific dimensions, platform format?
    • Mood: Authoritative, inviting, dramatic, clean?

    If the request is vague, ask one clarifying question about use case and style.

    Step 2: Select Domain Mode

    Choose the expertise lens for the image:

    Mode When to use Prompt emphasis
    Editorial Blog headers, feature images, lifestyle Styling, composition, publication references
    Product E-commerce posts, reviews, comparisons Surface materials, studio lighting, clean BG
    Landscape Environmental backgrounds, travel, hero sections Atmospheric perspective, depth layers, time of day
    UI/Web Tech blog icons, illustrations, diagrams Clean vectors, flat design, exact colors
    Infographic Data-driven posts, processes, comparisons Layout structure, hierarchy, accessible colors
    Abstract Pattern backgrounds, section dividers, decorative Color theory, mathematical forms, textures

    Load references/prompt-engineering-blog.md for domain mode modifier libraries.

    Step 3: Construct the 6-Component Reasoning Brief

    Build the prompt as natural narrative paragraphs, not keyword lists:

    1. Subject - Who/what, with rich physical detail (textures, materials, scale)
    2. Action - What is happening, pose, gesture, movement, state
    3. Context - Environment, setting, time of day, season, weather
    4. Composition - Camera angle, shot type, framing, negative space, depth
    5. Lighting - Light source, quality, direction, color temperature, shadows
    6. Style - Art medium, aesthetic, film stock, reference artists/eras

    Template for photorealistic blog images:

    A photorealistic [shot type] of [subject with physical detail], [action/pose],
    set in [environment with specifics]. [Lighting conditions] create [mood].
    Captured with [camera model], [focal length] lens at [f-stop], producing
    [depth of field effect]. [Color palette/grading notes]. Aspect ratio 16:9,
    suitable as a blog [hero image/inline illustration] at [target dimensions].
    

    Template for illustrated/stylized:

    A [art style] [format] of [subject with character detail], featuring
    [distinctive characteristics] with [color palette]. [Line style] and
    [shading technique]. Background is [description]. [Mood/atmosphere].
    

    Step 4: Set Aspect Ratio

    Call set_aspect_ratio BEFORE generating. Use conversation_id: "default".

    Blog Use Case Ratio
    Hero / Cover / OG 16:9
    Product shot / Square 4:3 or 1:1
    Section divider 21:9, then crop wider in post-processing if needed
    Vertical (stories) 9:16

    Step 5: Generate via MCP

    MCP Tool When
    set_aspect_ratio Always call first, even for 1:1
    gemini_generate_image New image from crafted prompt
    gemini_edit_image Modify existing image
    gemini_chat Iterative refinement / multi-turn sessions
    get_image_history Review generated images with conversation_id: "default"
    clear_conversation Reset session context

    Model selection:

    • Stable Google API IDs: gemini-3.1-flash-image and gemini-3-pro-image
    • Pinned @ycse/nanobanana-mcp@1.1.1: set_model accepts flash and pro, but maps them to preview IDs that shut down on 2026-06-25
    • Use direct API or a newer MCP package that explicitly supports stable image IDs before promising working MCP image generation

    Load references/mcp-tools.md for parameter details. Load references/gemini-models.md for model specs, pricing, and rate limits.

    Step 6: Post-Processing (when needed)

    After generation, resize/convert for blog use:

    # Resize to blog hero dimensions (1200x630)
    magick input.png -resize 1200x630^ -gravity center -extent 1200x630 hero.png
    
    # Convert to WebP for web optimization
    magick input.png -quality 85 output.webp
    
    # Convert to AVIF when target browsers support it
    magick input.png -quality 80 output.avif
    
    # Crop to exact OG dimensions
    magick input.png -resize 1200x630^ -gravity center -extent 1200x630 og-image.png
    

    Check if magick (ImageMagick 7) is available. Fall back to convert if not.

    Step 7: Deliver

    Provide:

    1. Image path - where it was saved (~/Documents/nanobanana_generated/)
    2. Crafted prompt - show the full Reasoning Brief (educational)
    3. Settings - model, aspect ratio, domain mode
    4. Alt text - descriptive sentence, 10-125 chars, topic keywords naturally
    5. Frontmatter snippet (for hero/OG images):
    coverImage: "/path/to/generated-image.png"
    coverImageAlt: "Descriptive alt text sentence with topic keywords"
    ogImage: "/path/to/generated-image.png"
    
    1. Refinement suggestions - 1-2 ideas if relevant

    Edit Workflow

    For /blog image edit <path> <instructions>:

    1. Read the image path and edit instruction
    2. Enhance the instruction (never pass raw):
      User says Claude crafts
      "remove background" Detailed edge-preserving background removal
      "make it warmer" Specific color temperature shift with preservation notes
      "add text" Font style, size, placement, contrast, readability notes
      "make it brighter" Increase exposure, lift shadows, maintain highlights
      "crop for social" Resize to 1200x630 with center-gravity crop
    3. Call gemini_edit_image with enhanced instruction
    4. Return modified image path and description

    Internal API (for blog-write / blog-rewrite)

    When invoked as a Task subagent from blog-write or blog-rewrite:

    Input (provided by calling skill):

    • image_type: hero, inline, og, divider
    • topic: blog post topic/title
    • section_context: (optional) heading or section the image supports
    • style_preference: (optional) photorealistic, illustrated, editorial
    • count: (optional) number of images needed (default: 1)

    Output (returned to calling skill):

    ### Generated Image
    - **Path:** ~/Documents/nanobanana_generated/image_timestamp.png
    - **Alt Text:** Descriptive sentence about the image
    - **Type:** hero / inline / og
    - **Domain Mode:** Editorial
    - **Aspect Ratio:** 16:9
    - **Suggested Frontmatter:**
      coverImage: "/path/to/image.png"
      coverImageAlt: "Alt text here"
    

    Graceful fallback: If MCP is unavailable, return immediately with no error. The calling workflow continues with stock photos. Never block blog-write or blog-rewrite because image generation is unavailable.

    Alt Text Generation

    For every generated image, create alt text following blog standards:

    • Full descriptive sentence (not keyword list)
    • 10-125 characters
    • Include topic keywords naturally
    • Describe what the image shows AND its relevance to the content
    • For charts/infographics: include the key data point

    Good: Marketing team analyzing AI search traffic data on a dashboard showing citation metrics Bad: SEO AI marketing blog optimization image

    Setup

    For /blog image setup:

    1. Run python3 skills/blog-image/scripts/setup_image_mcp.py (interactive)
      • Prefer: GOOGLE_AI_API_KEY=... python3 skills/blog-image/scripts/setup_image_mcp.py
      • Or: python3 skills/blog-image/scripts/setup_image_mcp.py --key-file /path/to/key.txt
      • Avoid --key unless necessary because command arguments can enter shell history and process lists
      • Default writes to ~/.claude/settings.json (user-private, mode 0600)
      • --project flag opts into project .mcp.json (env-expansion only, refuses to write a literal key into a tracked file)
    2. Verify: python3 skills/blog-image/scripts/validate_image_setup.py
    3. Requires:
    4. The script pins the package to @ycse/nanobanana-mcp@1.1.1. That npm release hard-codes preview image model IDs that shut down on 2026-06-25. Update setup, validation, and this documentation together when a package release with stable ID support is available.

    Safety Filter Auto-Rephrase

    When IMAGE_SAFETY or SAFETY is returned, do NOT give up. Auto-rephrase and retry:

    1. Identify the likely trigger (violence, public figures, NSFW-adjacent, or overly cautious filter)
    2. Rephrase using positive framing - describe what you WANT, not what to avoid
    3. If the subject is a person, make them generic (remove celebrity-like specifics)
    4. If the scene is dramatic, soften: "intense" → "focused", "battle" → "competition"
    5. Retry with the rephrased prompt (max 3 attempts before reporting to user)

    Google acknowledged filters "became way more cautious than we intended" - benign prompts are sometimes blocked. Persistence with rephrasing usually succeeds.

    Edit, Don't Re-roll

    If an image is 80% correct, use gemini_chat for conversational editing rather than regenerating from scratch. The session maintains style consistency, so targeted edits preserve what works while fixing what doesn't.

    When to edit vs regenerate:

    • Color slightly off → Edit ("shift the color temperature warmer")
    • Wrong composition entirely → Regenerate with revised brief
    • Good scene but wrong lighting → Edit ("change to golden hour lighting from the left")
    • Missing a detail → Edit ("add a steaming coffee cup on the desk")

    Error Handling

    Error Resolution
    MCP not configured Run /blog image setup
    API key invalid New key at https://aistudio.google.com/apikey
    Rate limited (429) Wait 60s, retry. Check live limits at https://ai.google.dev/gemini-api/docs/rate-limits
    IMAGE_SAFETY Auto-rephrase (see above) - Layer 2 filter, non-configurable
    PROHIBITED_CONTENT Content policy violation - topic is blocked. Non-retryable.
    SAFETY Rephrase prompt - Layer 1 filter
    Vague request Ask one clarifying question before generating
    Poor quality Review Reasoning Brief - likely missing lighting (biggest quality differentiator)
    MCP unavailable (internal call) Return silently - calling workflow uses stock photos

    Reference Documentation

    Load on-demand - do NOT load all at startup:

    • references/prompt-engineering-blog.md - Domain modes, 6-component system, blog templates
    • references/gemini-models.md - Model specs, rate limits, aspect ratios, pricing
    • references/mcp-tools.md - MCP tool parameters and response formats

    Reproducido de AgriciDaniel/claude-blog bajo licencia MIT. Leer esta página en markdown.

    Archivos

    6 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

    Antes de instalar

    Necesita Node.js 18+, el servidor `@ycse/nanobanana-mcp` configurado y una clave de Google AI; la versión fijada apunta a IDs de modelo en vista previa que se retiran el 2026-06-25.

    Variables de entorno:GOOGLE_AI_API_KEY

    Detalles

    Licencia
    MIT
    Recursos incluidos
    scripts en python + referencias
    Código fuente
    Ver SKILL.md

    Etiquetas

    Más de AgriciDaniel/claude-blog

    Este repo incluye 33 skills. Si instalas uno, normalmente ya tienes los demás. Ver el pack claude-blog entero y su comando de instalación

    Blog

    2.1k

    Motor de blog de ciclo completo con 31 subskills, 12 plantillas, puntuación sobre 100 y 5 agentes. Enruta cada petición a la subskill correcta: escribir, reescribir, analizar, auditar, schema, clusters y publicación multilingüe.

    Costo de contexto al activarse
    6.2k tok
    Tamaño del paquete
    35 archivos
    Última actualización
    hace 19 días
    seo geo

    Integración con las APIs de Google para rendimiento de blog: PageSpeed Insights, CrUX con 25 semanas de histórico, Search Console, URL Inspection, Indexing API, GA4, NLP de entidades, YouTube y Keyword Planner.

    Costo de contexto al activarse
    3.3k tok
    Tamaño del paquete
    24 archivos
    Última actualización
    hace 19 días
    seo geo

    Consulta cuadernos de Google NotebookLM para obtener respuestas ancladas en tus propios documentos y con citas: gestiona la biblioteca de cuadernos, la autenticación con Google y el descubrimiento de contenido.

    Costo de contexto al activarse
    2.5k tok
    Tamaño del paquete
    15 archivos
    Última actualización
    hace 19 días
    investigacion

    Genera narración en audio de posts con Google Gemini TTS: resumen hablado, lectura completa o diálogo tipo pódcast a dos voces, con 30 voces y salida MP3 más el código de inserción HTML5.

    Costo de contexto al activarse
    2.2k tok
    Tamaño del paquete
    8 archivos
    Última actualización
    hace 19 días
    redaccion contenido

    Motor de clusters temáticos semánticos: investiga keywords desde el SERP, agrupa por intención y solapamiento, construye una arquitectura hub-and-spoke, genera un mapa SVG y ejecuta el cluster llamando a blog-write.

    Costo de contexto al activarse
    4.9k tok
    Tamaño del paquete
    4 archivos
    Última actualización
    hace 19 días
    seo geo

    Integra el marco FLOW (Find, Optimize, Win) para blogs: ejecuta los prompts de cada etapa desde una base de 30 prompts aplicables a blog, con licencia CC BY 4.0 y sincronización desde el repositorio de FLOW.

    Costo de contexto al activarse
    2k tok
    Tamaño del paquete
    36 archivos
    Última actualización
    hace 19 días
    seo geo

    Skills relacionados

    Audita y puntúa posts con 100 puntos en 5 categorías: calidad de contenido, SEO, señales E-E-A-T, elementos técnicos y preparación para citas de IA. Exporta en markdown, JSON o tabla y admite análisis por lotes.

    Costo de contexto al activarse
    3.8k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    el mes pasado
    redaccion contenido

    Genera narración en audio de posts con Google Gemini TTS: resumen hablado, lectura completa o diálogo tipo pódcast a dos voces, con 30 voces y salida MP3 más el código de inserción HTML5.

    Costo de contexto al activarse
    2.2k tok
    Tamaño del paquete
    8 archivos
    Última actualización
    hace 19 días
    redaccion contenido

    Genera calendarios editoriales con clusters temáticos, cadencia de publicación, revisiones por cambio material, oportunidades estacionales, fórmula de mezcla de contenidos y planificación de distribución, mensual o trimestral.

    Costo de contexto al activarse
    3k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    el mes pasado
    redaccion contenido