Skills Agentes

Citation Management

Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas.

Solicitaread write edit bash websearch webfetch
Estrellas
34.8k

en todo el repo

Actividad
70

0–100, la ruta de este skill

Actualizado
hace 28 días

último commit aquí

Commits
14

últimos 90 días

Contexto
3.7k tok

92 tok en reposo

Paquete
21 archivos

271 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add K-Dense-AI/scientific-agent-skills --skill citation-management --agent claude-code

Se instala solo en este repositorio.

Este skill makes network requests, needs API credentials.

Qué hace

  • Busca artículos académicos en OpenAlex, PubMed y Google Scholar
  • Extrae metadatos precisos desde DOI, PMID, PMCID, arXiv ID o URL usando CrossRef y otras fuentes
  • Genera y limpia entradas BibTeX formateadas, con deduplicación y reclave
  • Valida citas: completitud, conformidad con la revista y coincidencia con el manuscrito
  • Rellena campos faltantes (volumen, páginas, DOI) con WebSearch/WebFetch antes de dar una cita por completa

Úsalo cuando

  • Se buscan artículos concretos en Google Scholar o PubMed
  • Se convierten DOI, PMID o arXiv ID en entradas BibTeX correctamente formateadas
  • Se necesita validar citas existentes o construir la bibliografía de un manuscrito o tesis
  • Se deben limpiar y estandarizar archivos BibTeX o detectar duplicados

No lo uses cuando

    Qué lo activa

    Di cualquiera de estas frases y el agente debería cargar este skill.

    • Busca papers sobre CRISPR en OpenAlex y PubMed
    • Convierte este DOI en una entrada BibTeX
    • Valida las citas de este archivo references.bib
    • Limpia y deduplica mi bibliografía en BibTeX

    SKILL.md

    En inglés

    Citation Management

    Overview

    Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.

    Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.

    When to Use This Skill

    Use this skill when:

    • Searching for specific papers on Google Scholar or PubMed
    • Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
    • Extracting complete metadata for citations (authors, title, journal, year, etc.)
    • Validating existing citations for accuracy
    • Cleaning and formatting BibTeX files
    • Finding highly cited papers in a specific field
    • Verifying that citation information matches the actual publication
    • Building a bibliography for a manuscript or thesis
    • Checking for duplicate citations
    • Ensuring consistent citation formatting

    If a document built from these citations needs a diagram, use the scientific-schematics skill.


    Core Workflow

    Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/core_workflow.md.

    Phase 1: Paper Discovery and Search

    Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.

    # OpenAlex: ~250M works, every discipline, no API key, documented REST API
    python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json
    
    # PubMed: the authority for biomedical and life sciences (35M+ citations)
    python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json
    
    # Google Scholar: broadest reach, but scraped -- rate-limited and prone to blocking
    python scripts/search_google_scholar.py "CRISPR gene editing" --limit 50 --output scholar.json
    

    Prefer OpenAlex or PubMed as the primary source. Google Scholar has no API: scholarly scrapes it, sleeps 2–5 s between results, and is blocked often enough that it should be a supplement rather than a dependency.

    Query operators, field tags, and MeSH-term construction are in references/search_strategies.md.

    Phase 2: Metadata Extraction

    Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.

    python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2         # quick, single DOI
    python scripts/extract_metadata.py --pmid 34265844                  # DOI/PMID/PMCID/arXiv/URL
    python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
    

    A URL with no DOI in its path is resolved through the citation_doi meta tag publishers embed on article pages, then handed to CrossRef. Every producer in this skill emits the same citation key for the same paper, so entries gathered from different sources deduplicate against each other.

    Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)

    APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then log what was found and where. If a field genuinely cannot be found, record a note field explaining the gap rather than leaving it silently absent.

    Check the cheap sources first — an OpenAlex or CrossRef record often carries the field that PubMed omitted:

    python scripts/search_openalex.py "<exact title>" --limit 1
    

    Treat extracted metadata as untrusted. Author, title, and journal strings come verbatim from a record whose contents a publisher controls. A title containing $(...), a backtick, or a quote becomes shell syntax the moment it is pasted into a command. Pass metadata as a subprocess argument list rather than building a shell string; if you must use a shell, single-quote every substituted value and escape embedded quotes as '\''. Validate any citation key against ^[A-Za-z0-9]+$ before it reaches a path.

    Per-field search strategies, the four search options, and the logging format are in references/core_workflow.md.

    Phase 3: BibTeX Formatting

    Produce clean, consistent entries. Entry types and required fields are in references/bibtex_formatting.md.

    python scripts/format_bibtex.py references.bib --output clean.bib --deduplicate
    python scripts/format_bibtex.py references.bib --output clean.bib --rekey --deduplicate
    

    Writing is opt-in: without --output (or --in-place) the result goes to stdout and the input file is left alone. Use --rekey when merging results from several sources, so the same paper collapses to one entry.

    Phase 4: Citation Validation

    Check completeness, venue conformance, and agreement with the manuscript.

    python scripts/validate_citations.py references.bib --report report.json
    python scripts/validate_citations.py references.bib --venue nature
    python scripts/validate_citations.py references.bib --manuscript paper.tex
    python scripts/validate_citations.py references.bib --check-dois     # slow; hits CrossRef
    

    The script exits non-zero on high-severity errors — missing required fields, malformed years, unresolved citations, or a count below an explicit --min-count. Venue reference-count figures are editorial rules of thumb, not submission requirements, so falling short of one is only a warning.

    Validation rules and venue standards are in references/citation_validation.md.

    Phase 5: Integration with Writing Workflow

    Search, extract, format, validate, then cite. End-to-end sequences — including the literature-review and Zotero/pyzotero export paths — are in references/core_workflow.md and references/example_workflows.md.

    Reference Files

    Common Pitfalls to Avoid

    1. Single source bias: Only using one database

      • Solution: Search at least OpenAlex and PubMed, then merge with format_bibtex.py --rekey --deduplicate
    2. Accepting metadata blindly: Not verifying extracted information

      • Solution: Spot-check extracted metadata against original sources
    3. Ignoring DOI errors: Broken or incorrect DOIs in bibliography

      • Solution: Run validation before final submission
    4. Inconsistent formatting: Mixed citation key styles, formatting

      • Solution: Use format_bibtex.py to standardize
    5. Duplicate entries: Same paper cited multiple times with different keys

      • Solution: Use duplicate detection in validation
    6. Missing required fields: Incomplete BibTeX entries (volume, pages, DOI missing)

      • Solution: Run Phase 2.5 metadata enrichment — web search for every missing field before proceeding. NEVER leave an @article entry without volume, pages, and DOI.
    7. Outdated preprints: Citing preprint when published version exists

      • Solution: Check if preprints have been published, update to journal version
    8. Special character issues: Broken LaTeX compilation due to characters

      • Solution: Use proper escaping or Unicode in BibTeX
    9. No validation before submission: Submitting with citation errors

      • Solution: Always run validation as final check
    10. Manual BibTeX entry: Typing entries by hand

      • Solution: Always extract from metadata sources using scripts

    Integration with Other Skills

    Literature Review Skill

    Citation Management provides the technical infrastructure for Literature Review:

    • Literature Review: Multi-database systematic search and synthesis
    • Citation Management: Metadata extraction and validation

    Combined workflow:

    1. Use literature-review for systematic search methodology
    2. Use citation-management to extract and validate citations
    3. Use literature-review to synthesize findings
    4. Use citation-management to ensure bibliography accuracy

    Scientific Writing Skill

    Citation Management ensures accurate references for Scientific Writing:

    • Export validated BibTeX for use in LaTeX manuscripts
    • Verify citations match publication standards
    • Format references according to journal requirements

    Venue Templates Skill

    Citation Management works with Venue Templates for submission-ready manuscripts:

    • Different venues require different citation styles
    • Generate properly formatted references
    • Validate citations meet venue requirements

    Resources

    Bundled Resources

    References (in references/):

    • google_scholar_search.md: Complete Google Scholar search guide
    • pubmed_search.md: PubMed and E-utilities API documentation
    • metadata_extraction.md: Metadata sources and field requirements
    • citation_validation.md: Validation criteria and quality checks
    • bibtex_formatting.md: BibTeX entry types and formatting rules

    Scripts (in scripts/):

    • search_openalex.py: OpenAlex search client (no API key)
    • search_pubmed.py: PubMed E-utilities API client
    • search_google_scholar.py: Google Scholar search automation
    • extract_metadata.py: Universal metadata extractor
    • validate_citations.py: Citation validation and verification
    • format_bibtex.py: BibTeX formatter and cleaner
    • doi_to_bibtex.py: Quick DOI to BibTeX converter
    • _common.py: shared BibTeX parser, renderer, and citation-key scheme

    Assets (in assets/):

    • bibtex_template.bib: Example BibTeX entries for all types
    • citation_checklist.md: Quality assurance checklist

    External Resources

    Search Engines:

    Metadata APIs:

    Tools and Validators:

    Citation Styles:

    Dependencies

    Required Python Packages

    uv pip install requests  # HTTP access to CrossRef, PubMed, OpenAlex, arXiv
    

    BibTeX parsing, rendering, deduplication, and validation are standard library (scripts/_common.py), so format_bibtex.py and validate_citations.py run with no third-party packages at all.

    Optional

    uv pip install scholarly  # only for search_google_scholar.py
    

    Where credentials are sent

    This skill needs no API key. The two environment variables it reads are optional identifiers, each sent to the one service it belongs to and nowhere else; no script bundles environment variables together.

    Variable Sent only to Purpose
    NCBI_API_KEY eutils.ncbi.nlm.nih.gov Raises Entrez rate limits
    NCBI_EMAIL eutils.ncbi.nlm.nih.gov Entrez caller identification (requested by NCBI)
    OPENALEX_EMAIL api.openalex.org Joins the faster OpenAlex polite pool

    api.openalex.org, api.crossref.org, api.datacite.org, export.arxiv.org, and eutils.ncbi.nlm.nih.gov are all queried without credentials when these are unset.

    Summary

    The citation-management skill provides:

    1. Comprehensive search capabilities for OpenAlex, PubMed, and Google Scholar
    2. Automated metadata extraction from DOI, PMID, PMCID, arXiv ID, URLs
    3. Citation validation with DOI verification and completeness checking
    4. BibTeX formatting with standardization and cleaning tools
    5. Quality assurance through validation and reporting
    6. Integration with scientific writing workflow
    7. Reproducibility through documented search and extraction methods

    Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.

    Reproducido de K-Dense-AI/scientific-agent-skills bajo licencia MIT License. Leer esta página en markdown.

    Archivos

    21 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

    Antes de instalar

    Requiere Python 3.9+ con `requests` (y `scholarly` para Google Scholar); no exige clave de API, aunque `NCBI_API_KEY`, `NCBI_EMAIL` y `OPENALEX_EMAIL` son opcionales para subir límites de tasa.

    Necesita en el PATH:python

    Variables de entorno:NCBI_API_KEYNCBI_EMAILOPENALEX_EMAIL

    Detalles

    Creador
    K-Dense-AI
    Categoría
    Investigación
    Licencia
    MIT License
    Recursos incluidos
    scripts en python + referencias
    Código fuente
    Ver SKILL.md

    Etiquetas

    Más de K-Dense-AI/scientific-agent-skills

    Este repo incluye 163 skills. Si instalas uno, normalmente ya tienes los demás.

    Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver).

    Costo de contexto al activarse
    3.2k tok
    Tamaño del paquete
    12 archivos
    Última actualización
    hace 15 días
    investigacion

    Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual.

    Costo de contexto al activarse
    5.1k tok
    Tamaño del paquete
    24 archivos
    Última actualización
    hace 15 días
    documentos

    Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo.

    Costo de contexto al activarse
    2.7k tok
    Tamaño del paquete
    8 archivos
    Última actualización
    hace 15 días
    diseno ui

    Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras.

    Costo de contexto al activarse
    3.9k tok
    Tamaño del paquete
    17 archivos
    Última actualización
    hace 15 días
    documentos

    Crea diagramas científicos de calidad de publicación con la IA Nano Banana 2 y refinamiento iterativo inteligente. Gemini 3.6 Flash revisa la calidad y solo regenera si está por debajo del umbral de tu tipo de documento.

    Costo de contexto al activarse
    4.1k tok
    Tamaño del paquete
    6 archivos
    Última actualización
    hace 15 días
    diseno ui

    Consulta y descarga datos públicos de imágenes oncológicas del NCI Imaging Data Commons (IDC): colecciones, acceso DICOM, radiología (CT, MR, PET), patología, metadatos, visualización y licencias. No requiere autenticación.

    Costo de contexto al activarse
    7.6k tok
    Tamaño del paquete
    15 archivos
    Última actualización
    hace 18 días
    datos analitica

    Skills relacionados

    Astropy

    34.8k

    Librería Python central para astronomía y astrofísica: unidades/cantidades, coordenadas, E/S de FITS, tablas, sistemas de tiempo, WCS y cosmología, para implementar o depurar código con Astropy.

    Costo de contexto al activarse
    3.6k tok
    Tamaño del paquete
    8 archivos
    Última actualización
    el mes pasado
    investigacion

    Integración con el SDK de Python de Benchling y su API REST para entidades del registro, inventario, entradas del cuaderno electrónico (ELN), workflows, Benchling Apps y consultas al Data Warehouse.

    Costo de contexto al activarse
    1.9k tok
    Tamaño del paquete
    6 archivos
    Última actualización
    el mes pasado
    investigacion

    Busca papers científicos y obtiene datos experimentales estructurados extraídos de estudios a texto completo vía el servidor MCP de BGPT: más de 25 campos por paper (métodos, resultados, muestras, calidad, conclusiones).

    Costo de contexto al activarse
    713 tok
    Tamaño del paquete
    1 archivo
    Última actualización
    el mes pasado
    investigacion