# Baoyu Url To Markdown > Obtiene cualquier URL y la convierte a markdown usando la CLI baoyu-fetch (Chrome CDP con adaptadores por sitio), con soporte para login/CAPTCHA. Fuente: https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-url-to-markdown Markdown: https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-url-to-markdown.md Repositorio: https://github.com/JimLiu/baoyu-skills Autor: JimLiu Licencia: MIT Actualizado: hace 3 meses Coste de contexto: 77 tok instalada, 2.1k tok al activarse, 69.6k tok con todos los archivos del bundle Bundle: 48 archivos, 272 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent claude-code # Cursor npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent cursor # Codex npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent codex # Gemini CLI npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent gemini # Windsurf npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent windsurf # Cline npx -y skills add JimLiu/baoyu-skills --skill baoyu-url-to-markdown --agent cline ``` ## Qué hace - Ejecuta la CLI baoyu-fetch (Chrome CDP) para descargar una URL y convertirla a markdown limpio - Aplica adaptadores específicos para X/Twitter, transcripciones de YouTube, hilos de Hacker News y páginas genéricas vía Defuddle - Gestiona login/CAPTCHA con modos de espera de interacción (--wait-for) - Descarga imágenes y videos localmente y reescribe los enlaces del markdown si se pide - Construye la ruta de salida y aplica un chequeo de calidad tras cada captura headless ## Cuándo usarla - El usuario quiere guardar una página web como markdown - Se necesita extraer una transcripción de YouTube, un hilo de X o de Hacker News - La página requiere login o CAPTCHA antes de poder leerse ## Qué la activa - "Guarda esta página como markdown" - "Convierte este hilo de X en un archivo markdown" - "Extrae la transcripción de este video de YouTube" - "Descarga este hilo de Hacker News con sus comentarios" ## Antes de instalar - Requiere el runtime bun y ejecutar `bun install` en scripts/ la primera vez; usa Chrome vía CDP. - Variables de entorno: DEFUDDLE_API_ORIGIN, HN_BASE_URL, JSON, READER - makes network requests - reads environment config ## Archivos - SKILL.md — 8 KB - references/adapters.md — 3 KB - references/config/first-time-setup.md — 2 KB - references/quality-gate.md — 2 KB - scripts/baoyu-fetch — 122 B - scripts/bun.lock — 34 KB - scripts/lib/adapters/generic/index.ts — 2 KB - scripts/lib/adapters/hn/index.ts — 10 KB - scripts/lib/adapters/index.ts — 834 B - scripts/lib/adapters/types.ts — 2 KB - scripts/lib/adapters/x/article.ts — 12 KB - scripts/lib/adapters/x/index.ts — 4 KB - scripts/lib/adapters/x/login.ts — 2 KB - scripts/lib/adapters/x/match.ts — 297 B - scripts/lib/adapters/x/payloads.ts — 2 KB - scripts/lib/adapters/x/session.ts — 1 KB - scripts/lib/adapters/x/shared.ts — 13 KB - scripts/lib/adapters/x/single.ts — 3 KB - scripts/lib/adapters/x/thread-loader.ts — 9 KB - scripts/lib/adapters/x/thread.ts — 9 KB - scripts/lib/adapters/x/types.ts — 600 B - scripts/lib/adapters/youtube/index.ts — 898 B - scripts/lib/adapters/youtube/transcript.ts — 13 KB - scripts/lib/adapters/youtube/utils.ts — 6 KB - scripts/lib/browser/cdp-client.ts — 7 KB - scripts/lib/browser/chrome-launcher.ts — 6 KB - scripts/lib/browser/cookie-sidecar.ts — 3 KB - scripts/lib/browser/interaction-gates.ts — 4 KB - scripts/lib/browser/network-journal.ts — 7 KB - scripts/lib/browser/page-snapshot.ts — 3 KB - scripts/lib/browser/profile.ts — 6 KB - scripts/lib/browser/session.ts — 5 KB - scripts/lib/cli.ts — 7 KB - scripts/lib/commands/convert.ts — 17 KB - scripts/lib/extract/document.ts — 846 B - scripts/lib/extract/html-cleaner.ts — 11 KB - scripts/lib/extract/html-extractor.ts — 2 KB - scripts/lib/extract/html-to-markdown.ts — 21 KB - scripts/lib/extract/markdown-renderer.ts — 5 KB - scripts/lib/media/default-downloader.ts — 5 KB - scripts/lib/media/markdown-media.ts — 13 KB - scripts/lib/media/media-utils.ts — 6 KB - scripts/lib/media/types.ts — 710 B - scripts/lib/types/defuddle-node.d.ts — 605 B - scripts/lib/types/shims.d.ts — 422 B - scripts/lib/utils/logger.ts — 645 B - scripts/lib/utils/url.ts — 300 B - scripts/package.json — 557 B ## SKILL.md Reproducido tal cual desde JimLiu/baoyu-skills bajo MIT. Esta sección es el documento original y está en inglés. # URL to Markdown Fetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown. ## User Input Tools When this skill prompts the user, follow this tool-selection rule (priority order): 1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent. 2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question. 3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order. Concrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes. ## CLI Setup **Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies. **Agent Execution Instructions**: 1. Determine this SKILL.md file's directory path as `{baseDir}` 2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun 3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts` 4. `${READER}` = `{baseDir}/scripts/baoyu-fetch` 5. Replace all `${READER}` in this document with the resolved value ## Preferences (EXTEND.md) Check EXTEND.md in priority order — the first one found wins: | Priority | Path | Scope | |----------|------|-------| | 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project | | 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG | | 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home | | Result | Action | |--------|--------| | Found | Read, parse, apply settings | | Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults | **EXTEND.md supports**: download media by default, default output directory. ### First-Time Setup ⛔ BLOCKING When EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call: - **Q1 — Media** (header "Media"): "How to handle images and videos in pages?" - "Ask each time (Recommended)" — Prompt after each save - "Always download" — Download to local `imgs/` and `videos/` - "Never download" — Keep remote URLs - **Q2 — Output** (header "Output"): "Default output directory?" - "url-to-markdown (Recommended)" — Save to `./url-to-markdown/{domain}/{slug}.md` - User may pick "Other" and type a custom path - **Q3 — Save** (header "Save"): "Where to save preferences?" - "User (Recommended)" — `~/.baoyu-skills/` (all projects) - "Project" — `.baoyu-skills/` (this project only) After answers, write EXTEND.md, confirm "Preferences saved to [path]", then continue. Full template: [references/config/first-time-setup.md](references/config/first-time-setup.md). ### Supported Keys | Key | Default | Values | Description | |-----|---------|--------|-------------| | `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always, `0` = never | | `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) | **EXTEND.md → CLI mapping**: | EXTEND.md key | CLI argument | Notes | |---------------|-------------|-------| | `download_media: 1` | `--download-media` | Requires `--output` to be set | | `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct flag | **Value priority**: CLI arguments → EXTEND.md → skill defaults. ## Usage ```bash # Default: headless capture, markdown to stdout ${READER} # Save to file ${READER} --output article.md # Save with media download ${READER} --output article.md --download-media # Wait for interaction (login/CAPTCHA) — auto-detect and continue ${READER} --wait-for interaction --output article.md # Wait for interaction — manual control (Enter to continue) ${READER} --wait-for force --output article.md # JSON output ${READER} --format json --output article.json # Force specific adapter ${READER} --adapter youtube --output transcript.md ``` ## Options | Option | Description | |--------|-------------| | `` | URL to fetch | | `--output ` | Output file path (default: stdout) | | `--format ` | Output format: `markdown` (default) or `json` | | `--json` | Shorthand for `--format json` | | `--adapter ` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) | | `--headless` | Force headless Chrome (no visible window) | | `--wait-for ` | Interaction wait mode: `none` (default), `interaction`, or `force` | | `--wait-for-interaction` | Alias for `--wait-for interaction` | | `--wait-for-login` | Alias for `--wait-for interaction` | | `--timeout ` | Page load timeout (default: 30000) | | `--interaction-timeout ` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) | | `--interaction-poll-interval ` | Poll interval for interaction checks (default: 1500) | | `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` | | `--media-dir ` | Base directory for downloaded media (default: same as `--output` directory) | | `--cdp-url ` | Reuse existing Chrome DevTools Protocol endpoint | | `--browser-path ` | Custom Chrome/Chromium binary path | | `--chrome-profile-dir ` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) | | `--debug-dir ` | Write debug artifacts (document.json, markdown.md, page.html, network.json) | ## Agent Quality Gate **CRITICAL**: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI. After every headless run, inspect the saved markdown. See [references/quality-gate.md](references/quality-gate.md) for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling. ## Output Path Generation The agent must construct the output file path — `baoyu-fetch` does not auto-generate paths. **Algorithm**: 1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/` 2. Extract domain from URL (e.g., `example.com`) 3. Generate slug from URL path or page title (kebab-case, 2-6 words) 4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated 5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md` Pass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained. ## Adapters & Media See [references/adapters.md](references/adapters.md) for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (`ask` / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts. ## Environment Variables | Variable | Description | |----------|-------------| | `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) | **Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA → `--wait-for interaction`. Debug → `--debug-dir` to inspect captured HTML and network logs. ## Extension Support Custom configurations via EXTEND.md. See **Preferences** section above for paths and supported keys. ## Dónde encaja - Categoría: [Herramientas para desarrolladores](https://skillsagentes.com/categorias/herramientas-desarrollo.md) — Skills que cambian cómo tu agente escribe, revisa y despliega código. - Creador: [JimLiu](https://skillsagentes.com/creators/jimliu.md) — 0 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Baoyu Image Gen](https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-image-gen.md): Generación de imágenes con IA usando OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate y Agnes; soporta texto a imagen, referencias, aspectos y lotes. - [Baoyu Post To Wechat](https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-post-to-wechat.md): Publica contenido en WeChat Official Account (微信公众号) vía API o Chrome CDP, con posts de artículo (文章) o de imagen-texto (贴图/图文). - [Baoyu Wechat Summary](https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-wechat-summary.md): Convierte los chats de un grupo de WeChat en un resumen estructurado usando el binario local wx-cli, con historial, perfiles de usuario y memoria de hechos por grupo entre ejecuciones. - [Baoyu Post To X](https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-post-to-x.md): Publica contenido y artículos en X (Twitter): posts normales con imágenes/vídeos y X Articles en Markdown, vía plugin Chrome de Codex, Computer Use o scripts CDP. - [Baoyu Cover Image](https://skillsagentes.com/skills/jimliu/baoyu-skills/baoyu-cover-image.md): Genera portadas de artículos con 5 dimensiones (tipo, paleta, renderizado, texto, mood), combinando 11 paletas y 7 estilos de renderizado, en formatos cinematic, widescreen o square. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)