# Gpt Image Edit > Edita imágenes con OpenAI GPT Image 2 (endpoint /edit) en RunComfy, con patrones de prompting documentados para preservación, texto multilingüe y multi-referencia (hasta 10 imágenes). Fuente: https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/gpt-image-edit Markdown: https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/gpt-image-edit.md Repositorio: https://github.com/prime-skills/runcomfy-agent-skills Autor: prime-skills Licencia: MIT Actualizado: hace 5 meses Coste de contexto: 175 tok instalada, 2.4k tok al activarse, 2.4k tok con todos los archivos del bundle Bundle: 1 archivo, 9 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent claude-code # Cursor npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent cursor # Codex npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent codex # Gemini CLI npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent gemini # Windsurf npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent windsurf # Cline npx -y skills add prime-skills/runcomfy-agent-skills --skill gpt-image-edit --agent cline ``` ## Qué hace - Edita imágenes con OpenAI GPT Image 2 (endpoint /edit) vía RunComfy, con patrones de prompting documentados para preservación de identidad y edición de texto multilingüe - Ejecuta `runcomfy run openai/gpt-image-2/edit` con hasta 10 imágenes de referencia - Sugiere routing hacia Nano Banana Edit, Flux Kontext o GPT Image 2 t2i según el caso de uso ## Cuándo usarla - Necesitas editar texto multilingüe o incrustado en una imagen - Quieres preservar identidad al hacer ediciones dirigidas (cara, marca, pose) - Necesitas ediciones con precisión de layout (mover titular, cambiar CTA) - Composición multi-referencia con hasta 10 imágenes ## Cuándo no - Necesitas consistencia por lote en muchas imágenes de SKU (mejor Nano Banana Edit) - Buscas fotorrealismo en retratos (Nano Banana Pro gana en esa comparación) - Quieres generar desde cero en vez de editar (usa el skill hermano gpt-image-2) ## Qué la activa - "Edita esta imagen con GPT Image 2 manteniendo la cara y la pose" - "Cambia el titular en japonés conservando el layout con gpt image edit" - "Compón el sujeto de la imagen 1 en la escena de la imagen 2 con chatgpt image edit" ## Antes de instalar - Requiere RunComfy CLI (`npm i -g @runcomfy/cli`), una cuenta RunComfy (`runcomfy login`) o la variable RUNCOMFY_TOKEN en CI/contenedores. - Necesita en el PATH: npx - makes network requests ## Archivos - SKILL.md — 9 KB ## SKILL.md Reproducido tal cual desde prime-skills/runcomfy-agent-skills bajo MIT. Esta sección es el documento original y está en inglés. # GPT Image Edit — Pro Pack on RunComfy [runcomfy.com](https://www.runcomfy.com/?utm_source=skills.sh&utm_medium=skill&utm_campaign=gpt-image-edit) · [Edit endpoint](https://www.runcomfy.com/models/openai/gpt-image-2/edit?utm_source=skills.sh&utm_medium=skill&utm_campaign=gpt-image-edit) · [Text-to-image sibling](https://www.runcomfy.com/models/openai/gpt-image-2/text-to-image?utm_source=skills.sh&utm_medium=skill&utm_campaign=gpt-image-edit) · [GitHub](https://github.com/agentspace-so/runcomfy-skills/tree/main/gpt-image-edit) OpenAI **GPT Image 2 — `/edit` endpoint** (ChatGPT Images 2.0 image-to-image) on the **RunComfy Model API**. Strongest in its class at preserving identity through targeted edits and rewriting embedded text in any script (Latin, kana, CJK, Cyrillic, Arabic). ```bash npx skills add agentspace-so/runcomfy-skills --skill gpt-image-edit -g ``` ## When to pick this model (vs siblings) | You want | Use | |---|---| | Edit multilingual / embedded text in image | **GPT Image Edit** | | Identity preservation through translated headline variants | **GPT Image Edit** | | Layout-precise edit (move headline, swap CTA, etc.) | **GPT Image Edit** | | Up to 10 reference images | **GPT Image Edit** | | Batch up to 20 images consistently | Nano Banana Edit | | Single-shot precise local edit, source-fidelity-first | Flux Kontext | | Generate from scratch with GPT Image 2 | sibling [`gpt-image-2`](../gpt-image-2) skill | | Batch SKU galleries with stable identity | Nano Banana Edit | ## Prerequisites 1. **RunComfy CLI** — `npm i -g @runcomfy/cli` 2. **RunComfy account** — `runcomfy login` opens a browser device-code flow. 3. **CI / containers** — set `RUNCOMFY_TOKEN=` instead of `runcomfy login`. ## Endpoints + input schema ### `openai/gpt-image-2/edit` | Field | Type | Required | Default | Notes | |---|---|---|---|---| | `prompt` | string | yes | — | Edit instruction. Lead with preservation, end with the change. | | `images` | string[] | yes | — | **Up to 10** publicly-fetchable HTTPS URLs. First is primary; rest are auxiliary. | | `size` | enum | no | `auto` | `auto` (preserve input), `1024_1024` (1:1), `1024_1536` (2:3 portrait), `1536_1024` (3:2 landscape). | `size=auto` preserves the input ratio — strongly recommended unless the edit explicitly changes framing. ## How to invoke **Single-ref preservation edit:** ```bash runcomfy run openai/gpt-image-2/edit \ --input '{ "prompt": "Keep the person'\''s face, pose, and brand mark unchanged. Replace the background with a soft warm-grey studio sweep and a gentle floor shadow.", "images": ["https://.../portrait.jpg"] }' \ --output-dir ``` **Multilingual text rewrite (preserve everything except the headline):** ```bash runcomfy run openai/gpt-image-2/edit \ --input '{ "prompt": "Keep the photograph, layout, and brand mark exactly as in the input. Replace only the in-image headline. The new headline reads \"今日のおすすめ\" in bold Japanese kana, same position and font weight as before.", "images": ["https://.../poster-en.jpg"] }' \ --output-dir ``` **Multi-ref composition:** ```bash runcomfy run openai/gpt-image-2/edit \ --input '{ "prompt": "Compose subject from image 1 into the room from image 2. Match the lighting and color palette of image 2. Keep image 1 subject identity (face, pose, clothing) unchanged.", "images": ["https://.../subject.jpg", "https://.../room.jpg"] }' \ --output-dir ``` ## Prompting — what actually works **Lead with preservation goals.** Always: `"Keep [face / pose / clothing / brand / framing] unchanged."` Then state the change. The model honors what's stated up front. **Multilingual text — quote the characters, name the script.** `"the headline reads \"コーヒー\" in bold Japanese kana"`, `"the label says \"АРОМА\" in Cyrillic, white on black"`, `"the right-margin caption reads \"تخفيض\" in Arabic right-to-left"`. Don't paraphrase — quote. **Directional language for spatial edits.** Concrete spatial scopes work: `"move the headline from top-right to bottom-center"`, `"remove the leftmost object only"`, `"replace the watermark in the bottom-right corner"`. **Multi-ref numbering.** When passing multiple `images`, refer to them by number: `"subject from image 1, lighting from image 2, color palette from image 3"`. The model routes cues correctly. **Use `size: "auto"` to preserve input ratio.** Only override when the edit explicitly changes framing (e.g. cropping a 16:9 to 1:1). **Anti-patterns:** - Long compound edit instructions ("change A and B and C and D") → drift increases per added scope. - Missing preservation goals → model subtly rewrites the face / brand / framing. - Paraphrasing in-image text instead of quoting it → text comes out different. - Asking for `size` outside the 3 fixed values + `auto` → 422. ## Where it shines | Use case | Why GPT Image Edit | |---|---| | **Multilingual ad localization** | One source asset → many language variants of the same headline | | **Brand-safe headline / CTA swaps** | Layout precision + preservation language hold the rest stable | | **Multi-ref composition (subject from one, scene from another)** | Numbered refs route cues correctly | | **Layout-precise repositioning** | Directional language ("top-right to bottom-center") honored | | **Identity preservation across signage edits** | Strongest in class for face / brand preservation through targeted edits | ## Sample prompts (verified to produce strong results) **Background swap with full preservation (page example):** ``` Turn the background into a bright minimal white-to-soft-gray studio sweep with gentle floor shadow; add a large headline in-image that reads "OPEN STUDIO" in a bold clean sans-serif, high contrast, centered; keep the main person or product, pose, and face identity unchanged ``` **Multilingual variant:** ``` Keep the photograph, layout, lighting, and brand mark exactly as in the input. Replace only the in-image headline. The new headline reads "コーヒー" in bold Japanese kana, same position and font weight as before. ``` **Multi-ref composition:** ``` Compose subject from image 1 into the kitchen from image 2. Match the warm window light and color palette of image 2. Keep subject identity (face, pose, clothing) from image 1 unchanged. ``` ## Limitations - **`size`: 3 fixed values + `auto`** — anything else 422s. - **`images`: up to 10** — first is primary, rest are auxiliary cues. - **Long compound prompts drift** — split into multiple passes when needed. - **For batch consistency across many SKU images, Nano Banana Edit (up to 20) is better.** - **Photorealism on portraits** — Nano Banana Pro wins head-to-head. ## Exit codes | code | meaning | |---|---| | 0 | success | | 64 | bad CLI args | | 65 | bad input JSON / schema mismatch | | 69 | upstream 5xx | | 75 | retryable: timeout / 429 | | 77 | not signed in or token rejected | Full reference: [docs.runcomfy.com/cli/troubleshooting](https://docs.runcomfy.com/cli/troubleshooting?utm_source=skills.sh&utm_medium=skill&utm_campaign=gpt-image-edit). ## How it works The skill invokes `runcomfy run openai/gpt-image-2/edit` with a JSON body matching the schema. The CLI POSTs to `https://model-api.runcomfy.net/v1/models/openai/gpt-image-2/edit`, polls the request, fetches the result, and downloads any `.runcomfy.net`/`.runcomfy.com` URL into `--output-dir`. `Ctrl-C` cancels the remote request before exit. ## Security & Privacy - **Token storage**: `runcomfy login` writes the API token to `~/.config/runcomfy/token.json` with mode 0600 (owner-only read/write). Set `RUNCOMFY_TOKEN` env var to bypass the file entirely in CI / containers. - **Input boundary**: the user prompt is passed as a JSON string to the CLI via `--input`. The CLI does NOT shell-expand the prompt; it transmits the JSON body directly to the Model API over HTTPS. No shell injection surface from prompt content. - **Third-party content**: image / mask / video URLs you pass are fetched by the RunComfy model server, not by the CLI on your machine. Treat external URLs as untrusted; image-based prompt injection is a known risk for any image-edit / video-edit model. - **Outbound endpoints**: only `model-api.runcomfy.net` (request submission) and `*.runcomfy.net` / `*.runcomfy.com` (download whitelist for generated outputs). No telemetry, no callbacks. - **Generated-file size cap**: the CLI aborts any single download > 2 GiB to prevent disk-fill from a malicious or runaway model output. ## Dónde encaja - Categoría: [Diseño y UI](https://skillsagentes.com/categorias/diseno-ui.md) — Sistemas de diseño, trabajo con componentes y acabado visual. - Creador: [prime-skills](https://skillsagentes.com/creators/prime-skills.md) — 30 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Ace Step](https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/ace-step.md): Genera, inpaint y outpaint música con ACE Step de StepFun-AI en RunComfy vía la CLI `runcomfy`: composición por tags, letras multilingües, hasta 4 min, desde $0.0002/s. - [Ai Music](https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/ai-music.md): Genera música con IA en RunComfy mediante la CLI `runcomfy`, enrutando entre ElevenLabs AI Music Generation (voz premium 44.1 kHz) y ACE Step / ACE Step 1.5 (código abierto, mucho más barato), más inpaint y outpaint de audio. - [Ai Avatar Video](https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/ai-avatar-video.md): Crea videos de avatar IA, talking-head y lip-sync en RunComfy con el CLI `runcomfy`, eligiendo entre OmniHuman, Wan 2-7, HappyHorse 1.0 y Seedance v2 Pro según la intención del usuario. - [Ai Image Generation](https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/ai-image-generation.md): Genera y edita imágenes en RunComfy vía la CLI `runcomfy`: un router inteligente entre todo el catálogo de modelos de imagen (FLUX 2, Nano Banana, GPT Image 2, Seedream, Qwen, Wan) para t2i e i2i. - [Ai Video Generation](https://skillsagentes.com/skills/prime-skills/runcomfy-agent-skills/ai-video-generation.md): Genera videos con IA en RunComfy vía el CLI `runcomfy`: un enrutador inteligente sobre todo el catálogo de modelos de video (HappyHorse, Wan 2-7, Seedance, Kling, Veo 3-1, Hailuo, Dreamina) para text-to-video, image-to-video y extend-video. ## Skills relacionadas - [Algorithmic Art](https://skillsagentes.com/skills/anthropics/skills/algorithmic-art.md): Crea arte algorítmico con p5.js, aleatoriedad con semilla y exploración interactiva de parámetros. Úsalo cuando pidan arte por código, arte generativo, flow fields o sistemas de partículas. - [Brand Guidelines](https://skillsagentes.com/skills/anthropics/skills/brand-guidelines.md): Aplica los colores y la tipografía oficiales de la marca Anthropic a cualquier artefacto que pueda beneficiarse de su look-and-feel. Úsalo cuando apliquen colores de marca o estándares de diseño. - [Canvas Design](https://skillsagentes.com/skills/anthropics/skills/canvas-design.md): Crea arte visual en documentos .png y .pdf partiendo de una filosofía de diseño. Úsalo cuando pidan un póster, una pieza de arte, un diseño u otra pieza estática. - [Frontend Design](https://skillsagentes.com/skills/anthropics/skills/frontend-design.md): Guía para un diseño visual distintivo e intencional al crear o rediseñar una UI: ayuda con dirección estética, tipografía y decisiones que no parezcan defaults templados. - [Slack Gif Creator](https://skillsagentes.com/skills/anthropics/skills/slack-gif-creator.md): Conocimiento y utilidades para crear GIFs animados optimizados para Slack: restricciones, herramientas de validación y conceptos de animación. Úsalo cuando pidan un GIF animado para Slack. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)