Skills Agentes

Genomic Coordinates

Convierte intervalos genómicos entre convenciones de coordenadas, normaliza y compara variantes, y detecta desajustes de assembly o nomenclatura de contigs antes de que corrompan un análisis.

Solicitaread write edit bash
Estrellas
34.8k

en todo el repo

Actividad
56

0–100, la ruta de este skill

Actualizado
el mes pasado

último commit aquí

Commits
1

últimos 90 días

Contexto
2.3k tok

230 tok en reposo

Paquete
10 archivos

101 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add K-Dense-AI/scientific-agent-skills --skill genomic-coordinates --agent claude-code

Se instala solo en este repositorio.

Qué hace

  • Convierte intervalos genómicos entre convenciones de coordenadas (BED, GFF/GTF, VCF, SAM/BAM, WIG, PSL, genePred, interval_list)
  • Normaliza y compara representaciones de variantes: recorta a parsimonia y alinea a la izquierda contra la referencia
  • Detecta discrepancias de assembly o de nomenclatura de contigs (chr1 vs 1, GRCh37 vs hg19 vs GRCh38 vs T2T)
  • Audita archivos BED/GTF/VCF en busca de violaciones de convención de coordenadas, con código de salida apto para CI
  • Mapea posiciones genómicas a transcrito, CDS o proteína, señalando los casos límite de HGVS

Úsalo cuando

  • Una coordenada cruza un límite de formato, herramienta o assembly y necesitas convertirla correctamente
  • Necesitas normalizar o comparar variantes escritas de formas distintas antes de unirlas o deduplicarlas
  • Sospechas un desajuste de assembly o de prefijo "chr" entre dos archivos que vas a cruzar
  • Necesitas auditar un BED, GTF o VCF en busca de errores de convención antes de usarlo

No lo uses cuando

    Qué lo activa

    Di cualquiera de estas frases y el agente debería cargar este skill.

    • Convierte estas coordenadas de BED a GFF
    • Normaliza esta variante y compárala con esta otra
    • Comprueba si este VCF y esta anotación usan el mismo assembly
    • Audita este archivo BED en busca de errores de coordenadas

    SKILL.md

    En inglés

    Genomic Coordinates

    When to use

    Any time a coordinate crosses a boundary: between two file formats, between two tools, between two assemblies, or between the genome and a transcript.

    The rule

    A coordinate is three facts, not one: the number, the convention it is written in, and the assembly it was measured against. Carry all three or the number is not interpretable.

    Coordinate errors are the quietest class of bug in genomics. An off-by-one BED file parses, sorts, and intersects without complaint. A GRCh37 VCF joined against a GRCh38 annotation returns rows. A right-shifted indel simply fails to match its entry in ClinVar, and the result is a variant reported as novel. Nothing raises an error; the answer is just wrong, and it is wrong in a direction that looks plausible.

    So: convert with the table, not from memory, and verify against the reference whenever a reference is available.

    The two conversions

    1-based inclusive  ->  0-based half-open :  start - 1,  end
    0-based half-open  ->  1-based inclusive :  start + 1,  end
    

    The end coordinate never moves. If a conversion changed both numbers, it is wrong.

    Which format is which

    0-based, half-open 1-based, inclusive
    BED, bedGraph, bigWig, narrowPeak GFF3, GTF, VCF
    BAM/CRAM (binary POS) SAM (text POS)
    PSL, genePred, refFlat WIG, Picard interval_list
    MAF (UCSC multiple alignment) MAF (TCGA mutation annotation)
    PyRanges, pybedtools GRanges/IRanges, samtools & UCSC & Ensembl region strings

    Both "MAF" formats exist, they mean different things, and they disagree. UCSC serves 0-based files through a 1-based browser box. references/format-conventions.md has the full table with per-format detail.

    cd skills/genomic-coordinates/scripts
    
    python3 convert_coords.py --list                          # the table
    python3 convert_coords.py --from bed --to gff chr1 999 1000
    python3 convert_coords.py --from ucsc --to bed "chr7:5,530,601-5,530,625"
    python3 convert_coords.py --from granges --to pyranges --input regions.tsv
    
    contig  input                 output           length  status  detail
    chr7    chr7:5530601-5530625  5530600-5530625  25      ok
    

    Zero-length BED features (chromStart == chromEnd, a legal insertion point) are reported as unrepresentable rather than converted to end = start - 1. Exit code is 1 when any interval is degenerate or invalid.

    Variants are not intervals

    A VCF POS for an indel is the anchor base — the base before the event, itself unchanged. And the same change can be written many ways: chr1:7:CAC:C, chr1:3:CAC:C and chr1:2:GCA:G are one deletion. Joining, deduplicating, or looking up variants before normalising loses real matches silently, and it loses them preferentially in repeats, where indels concentrate.

    Normalise — trim to parsimony, then left-align against the reference — before any comparison:

    python3 normalize_variant.py --fasta ref.fa chr1 7 CAC C
    python3 normalize_variant.py --fasta ref.fa --split --input cohort.vcf
    python3 normalize_variant.py --fasta ref.fa --compare chr1:7:CAC:C chr1:2:GCA:G
    
    input         normalized    type      pos_shift  ref_check  changed
    chr1:7:CAC:C  chr1:2:GCA:G  deletion  5          ok         yes
    

    Every record's REF is checked against the FASTA first. A MISMATCH means the variants and the reference are different assemblies — stop and run check_contigs.py rather than adjusting coordinates. Multi-allelic records must be split with --split before normalising, never after.

    HGVS shifts indels the opposite way, 3'-most along the transcript. For a minus-strand gene that is the opposite genomic direction from VCF's left-alignment. Details and the full procedure: references/variant-representation.md.

    Check the assembly before trusting a join

    python3 check_contigs.py --identify unknown.fa.fai
    python3 check_contigs.py variants.vcf annotation.gtf --genome GRCh38.fa.fai
    
    file          kind    contigs  naming        assembly  detail
    ref.fa.fai    sizes   25       plain         GRCh37    24/24 primary chromosome lengths match;
                                                           chrM is 16569 bp, i.e. GRCh37/38 (rCRS MT)
    

    The script reads .fai, .chrom.sizes, VCF headers, SAM headers, FASTA, BED, and GTF/GFF, identifies the assembly from primary-chromosome lengths, and reports every reason a join between two files would go wrong: naming mismatch, length conflict, coordinates past a contig end, contigs present in one file only. Exit code 1 on any incompatibility.

    GRCh37 and hg19 differ only in the mitochondrion — 16,569 bp (rCRS) versus 16,571 bp. Nuclear coordinates are identical, so a mixed pipeline runs fine and only the mtDNA results are wrong. check_contigs.py reports which one it found. Builds, naming schemes, ALT contigs, and liftover pitfalls: references/reference-builds.md.

    Audit a file against its own format

    python3 audit_intervals.py peaks.bed
    python3 audit_intervals.py gencode.gtf --genome hg38.chrom.sizes
    python3 audit_intervals.py cohort.vcf --genome GRCh38.fa.fai
    

    Looks for the evidence that a coordinate mistake leaves behind:

    Finding What it proves
    start_below_one in GFF/GTF 0-based data in a 1-based file; everything is one base left
    many_zero_length in BED 1-based single-base features written into a 0-based file
    past_contig_end wrong assembly, or an off-by-one at the contig edge
    mixed_contig_naming any join will silently match one subset
    first_block_offset BED12 blockStarts written as absolute coordinates
    not_parsimonious untrimmed alleles; normalise before joining
    bad_alt_allele Ensembl/VEP - notation in a VCF, which has no anchor base

    Exit code 1 on any fatal finding, so it works as a CI gate on a data directory.

    Transcript, CDS, and protein positions

    c.742 and chr17:7,674,220 are both "position", and neither converts to the other by arithmetic. Transcript coordinates count spliced bases in transcription order — decreasing genomic coordinate on the minus strand — and c.1 is the A of the initiator ATG, not the start of the transcript.

    The rules that get mis-remembered: there is no c.0; 5' UTR positions are negative and 3' UTR positions take a *; GFF phase is the bases to remove to reach the next codon, not start % 3; and a c. description is meaningless without a versioned transcript accession, because the same variant numbers differently in each transcript. references/transcript-coordinates.md has the conversion procedure and the boundary cases.

    Do the conversion with a tool that holds the transcript model — VEP, bcftools csq, Mutalyzer, the hgvs package — not by hand.

    Reporting results

    State the assembly next to the coordinates, every time. chr7:5,530,601-5,530,625 is not a location; chr7:5,530,601-5,530,625 (GRCh38) is. Say which convention a coordinate column is in, in the column header or the file's documentation. When a conversion produced a result, say which direction it went.

    References

    • references/format-conventions.md — every format's convention, with per-format detail, BED12 block rules, region-string syntax, and tool behaviour.
    • references/variant-representation.md — VCF allele conventions, the normalisation algorithm, equivalence checking, multi-allelic splitting, and how HGVS disagrees with VCF.
    • references/reference-builds.md — build signatures, GRCh37 vs hg19, ALT contigs, naming schemes, and liftover failure modes.
    • references/transcript-coordinates.md — genomic ↔ transcript ↔ CDS ↔ protein, HGVS numbering, phase, and transcript choice.

    Reproducido de K-Dense-AI/scientific-agent-skills bajo licencia MIT. Leer esta página en markdown.

    Archivos

    10 archivos en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

    Antes de instalar

    Requiere Python 3.11+ (solo librería estándar); normalizar variantes necesita un FASTA de referencia y usa su índice .fai si está disponible.

    Necesita en el PATH:python3

    Detalles

    Creador
    K-Dense-AI
    Licencia
    MIT
    Recursos incluidos
    scripts en python + referencias
    Código fuente
    Ver SKILL.md

    Etiquetas

    Más de K-Dense-AI/scientific-agent-skills

    Este repo incluye 163 skills. Si instalas uno, normalmente ya tienes los demás.

    Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas.

    Costo de contexto al activarse
    3.7k tok
    Tamaño del paquete
    21 archivos
    Última actualización
    hace 28 días
    investigacion

    Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver).

    Costo de contexto al activarse
    3.2k tok
    Tamaño del paquete
    12 archivos
    Última actualización
    hace 15 días
    investigacion

    Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual.

    Costo de contexto al activarse
    5.1k tok
    Tamaño del paquete
    24 archivos
    Última actualización
    hace 15 días
    documentos

    Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo.

    Costo de contexto al activarse
    2.7k tok
    Tamaño del paquete
    8 archivos
    Última actualización
    hace 15 días
    diseno ui

    Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras.

    Costo de contexto al activarse
    3.9k tok
    Tamaño del paquete
    17 archivos
    Última actualización
    hace 15 días
    documentos

    Crea diagramas científicos de calidad de publicación con la IA Nano Banana 2 y refinamiento iterativo inteligente. Gemini 3.6 Flash revisa la calidad y solo regenera si está por debajo del umbral de tu tipo de documento.

    Costo de contexto al activarse
    4.1k tok
    Tamaño del paquete
    6 archivos
    Última actualización
    hace 15 días
    diseno ui

    Skills relacionados

    Aeon

    34.8k

    Para tareas de machine learning con series temporales: clasificación, regresión, clustering, forecasting, detección de anomalías, segmentación y búsqueda de similitud, con APIs compatibles con scikit-learn.

    Costo de contexto al activarse
    3.1k tok
    Tamaño del paquete
    12 archivos
    Última actualización
    el mes pasado
    datos analitica

    Anndata

    34.8k

    Estructura de datos para matrices anotadas en análisis de célula única. Úsala con archivos .h5ad o el ecosistema scverse; para análisis usa scanpy, para modelos probabilísticos scvi-tools, para escala poblacional cellxgene-census.

    Costo de contexto al activarse
    3k tok
    Tamaño del paquete
    6 archivos
    Última actualización
    el mes pasado
    datos analitica

    Infiere redes de regulación génica (GRN) a partir de datos de expresión génica con algoritmos escalables (GRNBoost2, GENIE3), para transcriptómica bulk o de célula única, con computación distribuida.

    Costo de contexto al activarse
    2.1k tok
    Tamaño del paquete
    5 archivos
    Última actualización
    el mes pasado
    datos analitica