# Docx > Crea, lee, edita y manipula documentos Word (.docx) y plantillas (.dotx): informes, memos, cartas, tablas de contenido, imágenes, buscar-reemplazar, cambios controlados y comentarios. Fuente: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/docx Markdown: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/docx.md Repositorio: https://github.com/K-Dense-AI/scientific-agent-skills Autor: K-Dense-AI Licencia: Proprietary. LICENSE.txt has complete terms Actualizado: el mes pasado Coste de contexto: 209 tok instalada, 1.8k tok al activarse, 282.8k tok con todos los archivos del bundle Bundle: 61 archivos, 1.1 MB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent claude-code # Cursor npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent cursor # Codex npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent codex # Gemini CLI npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent gemini # Windsurf npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent windsurf # Cline npx -y skills add K-Dense-AI/scientific-agent-skills --skill docx --agent cline ``` ## Qué hace - Crea documentos Word nuevos con docx-js (npm), gestionando tamaño de página, tablas, listas, imágenes y saltos de página - Edita documentos .docx existentes descomprimiendo, modificando word/document.xml y recomprimiendo - Extrae el contenido de un .docx a Markdown con pandoc - Gestiona control de cambios (tracked changes) y comentarios con scripts dedicados y validación XSD - Renderiza el resultado a PDF/JPEG con LibreOffice y pdftoppm para verificarlo visualmente ## Cuándo usarla - El usuario pide crear un informe, memo, carta o plantilla como documento Word/.docx - Hay que extraer o reorganizar contenido de un archivo .docx o .dotx - Hay que insertar/reemplazar imágenes, hacer buscar-reemplazar o trabajar con cambios controlados o comentarios en Word - Convertir contenido en un documento Word bien formateado, con tabla de contenidos, encabezados o numeración de página ## Cuándo no - No usarlo para PDFs, hojas de cálculo, Google Docs ni tareas de código sin relación con generación de documentos ## Qué la activa - "Crea un informe en Word con tabla de contenidos y numeración de página" - "Añade comentarios a este contrato en .docx" - "Acepta todos los cambios controlados de este documento Word" - "Extrae el texto de este .docx a Markdown" ## Antes de instalar - Requiere el paquete npm docx (preinstalado), pandoc, LibreOffice (soffice) y pdftoppm (Poppler). - Necesita en el PATH: python - makes network requests ## Archivos - LICENSE.txt — 1 KB - SKILL.md — 7 KB - scripts/__init__.py — 1 B - scripts/accept_changes.py — 4 KB - scripts/comment.py — 14 KB - scripts/merge_runs.py — 9 KB - scripts/office/helpers/__init__.py — 3 KB - scripts/office/helpers/pptx_chart.py — 6 KB - scripts/office/helpers/pptx_slide.py — 2 KB - scripts/office/helpers/pptx_theme.py — 3 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-chart.xsd — 73 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-chartDrawing.xsd — 7 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-diagram.xsd — 50 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-lockedCanvas.xsd — 624 B - scripts/office/schemas/ISO-IEC29500-4_2016/dml-main.xsd — 148 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-picture.xsd — 1 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-spreadsheetDrawing.xsd — 9 KB - scripts/office/schemas/ISO-IEC29500-4_2016/dml-wordprocessingDrawing.xsd — 14 KB - scripts/office/schemas/ISO-IEC29500-4_2016/pml.xsd — 82 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-additionalCharacteristics.xsd — 1 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-bibliography.xsd — 7 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-commonSimpleTypes.xsd — 6 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-customXmlDataProperties.xsd — 1 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-customXmlSchemaProperties.xsd — 880 B - scripts/office/schemas/ISO-IEC29500-4_2016/shared-documentPropertiesCustom.xsd — 3 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-documentPropertiesExtended.xsd — 3 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-documentPropertiesVariantTypes.xsd — 7 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-math.xsd — 23 KB - scripts/office/schemas/ISO-IEC29500-4_2016/shared-relationshipReference.xsd — 1 KB - scripts/office/schemas/ISO-IEC29500-4_2016/sml.xsd — 237 KB - scripts/office/schemas/ISO-IEC29500-4_2016/vml-main.xsd — 26 KB - scripts/office/schemas/ISO-IEC29500-4_2016/vml-officeDrawing.xsd — 25 KB - scripts/office/schemas/ISO-IEC29500-4_2016/vml-presentationDrawing.xsd — 535 B - scripts/office/schemas/ISO-IEC29500-4_2016/vml-spreadsheetDrawing.xsd — 6 KB - scripts/office/schemas/ISO-IEC29500-4_2016/vml-wordprocessingDrawing.xsd — 4 KB - scripts/office/schemas/ISO-IEC29500-4_2016/wml.xsd — 167 KB - scripts/office/schemas/ISO-IEC29500-4_2016/xml.xsd — 5 KB - scripts/office/schemas/ecma/fouth-edition/opc-contentTypes.xsd — 2 KB - scripts/office/schemas/ecma/fouth-edition/opc-coreProperties.xsd — 2 KB - scripts/office/schemas/ecma/fouth-edition/opc-digSig.xsd — 3 KB - scripts/office/schemas/ecma/fouth-edition/opc-relationships.xsd — 1 KB - scripts/office/schemas/mce/mc.xsd — 3 KB - scripts/office/schemas/microsoft/wml-2010.xsd — 26 KB - scripts/office/schemas/microsoft/wml-2012.xsd — 4 KB - scripts/office/schemas/microsoft/wml-2018.xsd — 901 B - scripts/office/schemas/microsoft/wml-cex-2018.xsd — 2 KB - scripts/office/schemas/microsoft/wml-cid-2016.xsd — 1002 B - scripts/office/schemas/microsoft/wml-sdtdatahash-2020.xsd — 600 B - scripts/office/schemas/microsoft/wml-symex-2015.xsd — 745 B - scripts/office/soffice.py — 8 KB - scripts/office/validate.py — 6 KB - scripts/office/validators/__init__.py — 336 B - scripts/office/validators/base.py — 33 KB - scripts/office/validators/docx.py — 17 KB - scripts/office/validators/pptx.py — 16 KB - scripts/office/validators/redlining.py — 11 KB - scripts/templates/comments.xml — 3 KB - scripts/templates/commentsExtended.xml — 3 KB - scripts/templates/commentsExtensible.xml — 3 KB - scripts/templates/commentsIds.xml — 3 KB - scripts/templates/people.xml — 115 B ## SKILL.md Reproducido tal cual desde K-Dense-AI/scientific-agent-skills bajo Proprietary. LICENSE.txt has complete terms. Esta sección es el documento original y está en inglés. # DOCX creation, editing, and analysis A `.docx` is a ZIP archive of XML files. Choose your approach by task: | Task | Approach | |---|---| | **Create** a new document | Write a `docx` (npm) script — see gotchas below | | **Edit** an existing document | `unzip` → edit `word/document.xml` → `zip` (docx-js cannot open existing files) | | **Read** content | `pandoc -t markdown file.docx` | > Script paths below are relative to this skill's directory. ## Creating with docx-js — gotchas `docx` is preinstalled — do not run `npm install` first; write the script and `require('docx')` directly. Only if that require fails: `npm install docx`. The model knows the API; these are the footguns: - **Page size defaults to A4.** For US Letter set `page: { size: { width: 12240, height: 15840 } }` (DXA; 1440 = 1″). - **Landscape:** pass portrait dimensions and `orientation: PageOrientation.LANDSCAPE` — docx-js swaps width/height internally. - **Tables need dual widths:** set `columnWidths` on the table AND `width` on every cell, both in `WidthType.DXA` (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width. - **Table shading:** use `ShadingType.CLEAR`, never `SOLID` (renders black). - **Lists:** never insert `•` literally; use a `numbering` config with `LevelFormat.BULLET`. - **`ImageRun` requires `type:`** (`"png"`, `"jpg"`, …). - **`PageBreak` must be inside a `Paragraph`.** - **Never use `\n`** — use separate `Paragraph` elements. - **TOC:** headings must use built-in `HeadingLevel.*`; custom heading styles need `outlineLevel` set or they won't appear. - **Don't use a table as a horizontal rule** — use a paragraph bottom border instead. - **Dot-leader / right-aligned-on-same-line:** use `PositionalTab` (`alignment: PositionalTabAlignment.RIGHT`, `leader: PositionalTabLeader.DOT`) inside a `TextRun`, not literal `.` or space padding. ## Verify the output After writing a `.docx`, render it and look at it: ```bash python scripts/office/soffice.py --headless --convert-to pdf output.docx pdftoppm -jpeg -r 100 output.pdf page ls page-*.jpg # then Read the images ``` `pdftoppm` zero-pads page numbers to the width of the page count (`page-01.jpg`…`page-12.jpg`). ## Editing existing documents Legacy `.doc` files must be converted first: `python scripts/office/soffice.py --headless --convert-to docx file.doc`. ```bash unzip -q doc.docx -d unpacked/ find unpacked -type l -delete # strip symlink entries — docx from external parties is untrusted python scripts/merge_runs.py unpacked/ # coalesce fragmented runs so text is findable # edit unpacked/word/document.xml in place — do NOT reformat or pretty-print (cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .) python scripts/office/validate.py out.docx --original doc.docx # XSD checks; --auto-repair fixes common issues # redlining? add --author "" to check every edit is tracked ``` Word splits text across many `` runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. `merge_runs.py` merges adjacent identically-formatted runs in `word/document.xml` without changing content or rendering; it also accepts a `.docx` directly (`python scripts/merge_runs.py doc.docx -o merged.docx`). **Tracked changes:** when redlining, validate with `--author ""` (needs `--original`) — it reports any text you changed without a ``/`` around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in ``/`` with `w:id`, `w:author`, `w:date` attributes. Inside ``, the text element is ``, not ``. A deleted paragraph mark (``) means "merge this paragraph into the next" — so deleting a paragraph outright is that plus a `` around every run. The `` must come before the rPr's other children; their order is schema-enforced. To produce a clean copy with all tracked changes accepted: `python scripts/accept_changes.py in.docx out.docx`. Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are *all* deleted vanishes. Word does this; `accept_changes.py` and `pandoc --track-changes=accept` don't always. Both fail the same way — they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered: - `pandoc --track-changes=accept` never joins the paragraphs. - `accept_changes.py` (LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph. An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML. ## Comments Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing `document.xml` (saves an unzip/rezip cycle), `.docx`-direct mode otherwise: ```bash # Against an already-unpacked directory (preferred when also placing markers) python scripts/comment.py unpacked/ "Fees & expenses cap is too low" python scripts/comment.py unpacked/ "Agreed" --parent 0 # Against a .docx directly python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx ``` The script writes `comments.xml`, `commentsExtended.xml`, `commentsIds.xml`, `commentsExtensible.xml`, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the ``/``/`` snippet to add to `word/document.xml` so the comment anchors to specific text — until you place those markers, the comment exists but is not visible. ## Dependencies `docx` (npm, preinstalled — install only if `require('docx')` fails) · `pandoc` · LibreOffice (`soffice`) · `pdftoppm` (Poppler) --- *This skill is created and maintained by [Anthropic](https://github.com/anthropics/skills/tree/main/skills/docx). Vendored here unmodified except for frontmatter metadata; see LICENSE.txt for terms.* ## Dónde encaja - Categoría: [Documentos](https://skillsagentes.com/categorias/documentos.md) — Lee, escribe y transforma archivos PDF, DOCX, XLSX y PPTX. - Creador: [K-Dense-AI](https://skillsagentes.com/creators/k-dense-ai.md) — 163 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Citation Management](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/citation-management.md): Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas. - [Scientific Slides](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scientific-slides.md): Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual. - [Literature Review](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/literature-review.md): Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver). - [Infographics](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/infographics.md): Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo. - [Latex Posters](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/latex-posters.md): Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)