# Biopython > Kit de herramientas de biología molecular: manipulación de secuencias, parseo de archivos (FASTA/GenBank/PDB), filogenética y acceso a NCBI/PubMed (Bio.Entrez); para búsquedas rápidas usa gget, para varios servicios usa bioservices. Fuente: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/biopython Markdown: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/biopython.md Repositorio: https://github.com/K-Dense-AI/scientific-agent-skills Autor: K-Dense-AI Licencia: Biopython License Agreement Actualizado: el mes pasado Coste de contexto: 81 tok instalada, 4k tok al activarse, 25k tok con todos los archivos del bundle Bundle: 8 archivos, 98 KB Permisos que pide: read write edit bash ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent claude-code # Cursor npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent cursor # Codex npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent codex # Gemini CLI npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent gemini # Windsurf npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent windsurf # Cline npx -y skills add K-Dense-AI/scientific-agent-skills --skill biopython --agent cline ``` ## Qué hace - Manipula secuencias biológicas (ADN, ARN, proteína) y lee/escribe formatos de archivo (FASTA, GenBank, FASTQ, PDB, mmCIF) - Accede programáticamente a bases de datos de NCBI (GenBank, PubMed, Protein, Gene) vía Bio.Entrez - Ejecuta búsquedas BLAST (locales o vía NCBI web) y parsea sus resultados - Analiza estructuras de proteínas desde archivos PDB (distancias, ángulos, superposición, RMSD) - Crea, manipula y visualiza árboles filogenéticos, y calcula estadísticas de secuencia ## Cuándo usarla - Manipular secuencias biológicas o convertir entre formatos de archivo bioinformáticos - Acceder a bases de datos de NCBI (GenBank, PubMed) o ejecutar/parsear búsquedas BLAST - Analizar estructuras de proteínas desde archivos PDB o mmCIF - Crear o manipular árboles filogenéticos, o hacer alineamientos de secuencias ## Cuándo no - Para búsquedas rápidas y puntuales (usa gget en su lugar) - Para integrar múltiples servicios bioinformáticos a la vez (usa bioservices en su lugar) ## Qué la activa - "Lee este archivo FASTA y cuenta la longitud de cada secuencia" - "Ejecuta un BLAST de esta secuencia de ADN contra la base nt" - "Calcula la distancia entre estos dos carbonos alfa en esta estructura PDB" - "Construye un árbol filogenético a partir de este alineamiento" ## Antes de instalar - Requiere Python 3.10+, NumPy y Biopython (recomendado 1.87+); los ejemplos de Entrez y BLAST web necesitan red, y BLAST/MUSCLE locales necesitan esas herramientas instaladas aparte. - Necesita en el PATH: rg - Variables de entorno: NCBI_API_KEY - makes network requests - needs API credentials ## Archivos - SKILL.md — 16 KB - references/advanced.md — 14 KB - references/alignment.md — 10 KB - references/blast.md — 13 KB - references/databases.md — 12 KB - references/phylogenetics.md — 14 KB - references/sequence_io.md — 8 KB - references/structure.md — 13 KB ## SKILL.md Reproducido tal cual desde K-Dense-AI/scientific-agent-skills bajo Biopython License Agreement. Esta sección es el documento original y está en inglés. # Biopython: Computational Molecular Biology in Python ## Overview Biopython is a comprehensive set of freely available Python tools for biological computation. It provides functionality for sequence manipulation, file I/O, database access, structural bioinformatics, phylogenetics, and many other bioinformatics tasks. The current version is **Biopython 1.87** (released 30 March 2026). It supports **Python 3.10-3.14** and PyPy3.10, and requires NumPy. Biopython 1.87 also addresses **CVE-2025-68463** in `Bio.Entrez.Parser` when parsing untrusted files, so prefer 1.87+ for workflows that parse externally supplied Entrez XML. ## When to Use This Skill Use this skill when: - Working with biological sequences (DNA, RNA, or protein) - Reading, writing, or converting biological file formats (FASTA, GenBank, FASTQ, PDB, mmCIF, etc.) - Accessing NCBI databases (GenBank, PubMed, Protein, Gene, etc.) via Entrez - Running BLAST searches or parsing BLAST results - Performing sequence alignments (pairwise or multiple sequence alignments) - Analyzing protein structures from PDB files - Creating, manipulating, or visualizing phylogenetic trees - Finding sequence motifs or analyzing motif patterns - Calculating sequence statistics (GC content, molecular weight, melting temperature, etc.) - Performing structural bioinformatics tasks - Working with population genetics data - Any other computational molecular biology task ## Core Capabilities Biopython is organized into modular sub-packages, each addressing specific bioinformatics domains: 1. **Sequence Handling** - Bio.Seq and Bio.SeqIO for sequence manipulation and file I/O 2. **Alignment Analysis** - Bio.Align and Bio.AlignIO for pairwise and multiple sequence alignments 3. **Database Access** - Bio.Entrez for programmatic access to NCBI databases 4. **BLAST Operations** - Bio.Blast for running and parsing BLAST searches 5. **Structural Bioinformatics** - Bio.PDB for working with 3D protein structures 6. **Phylogenetics** - Bio.Phylo for phylogenetic tree manipulation and visualization 7. **Advanced Features** - Motifs, population genetics, sequence utilities, and more ## Installation and Setup Install the current stable Biopython release with an explicit version pin for reproducibility: ```bash uv pip install "biopython==1.87" ``` For NCBI database access, always set your email address (required by NCBI). For reusable software, set a stable `Entrez.tool` value and register the tool/email with NCBI. For higher rate limits (10 req/s instead of 3 req/s), read only `NCBI_API_KEY` from the environment — do not hardcode keys or load unrelated environment variables: ```python import os from Bio import Entrez Entrez.email = "your.email@example.com" # required — use your real email Entrez.tool = "your_tool_name" # optional but recommended for reusable software # Optional: register at https://www.ncbi.nlm.nih.gov/account/settings/ if api_key := os.environ.get("NCBI_API_KEY"): Entrez.api_key = api_key ``` ## Using This Skill This skill provides comprehensive documentation organized by functionality area. When working on a task, consult the relevant reference documentation: ### 1. Sequence Handling (Bio.Seq & Bio.SeqIO) **Reference:** `references/sequence_io.md` Use for: - Creating and manipulating biological sequences - Reading and writing sequence files (FASTA, GenBank, FASTQ, etc.) - Converting between file formats - Extracting sequences from large files - Sequence translation, transcription, and reverse complement - Working with SeqRecord objects **Quick example:** ```python from Bio import SeqIO # Read sequences from FASTA file for record in SeqIO.parse("sequences.fasta", "fasta"): print(f"{record.id}: {len(record.seq)} bp") # Convert GenBank to FASTA SeqIO.convert("input.gb", "genbank", "output.fasta", "fasta") ``` ### 2. Alignment Analysis (Bio.Align & Bio.AlignIO) **Reference:** `references/alignment.md` Use for: - Pairwise sequence alignment (global and local) - Reading and writing multiple sequence alignments - Using substitution matrices (BLOSUM, PAM) - Calculating alignment statistics - Customizing alignment parameters **Quick example:** ```python from Bio import Align # Pairwise alignment aligner = Align.PairwiseAligner() aligner.mode = 'global' alignments = aligner.align("ACCGGT", "ACGGT") print(alignments[0]) ``` ### 3. Database Access (Bio.Entrez) **Reference:** `references/databases.md` Use for: - Searching NCBI databases (PubMed, GenBank, Protein, Gene, etc.) - Downloading sequences and records - Fetching publication information - Finding related records across databases - Batch downloading with proper rate limiting **Quick example:** ```python from Bio import Entrez Entrez.email = "your.email@example.com" # Search PubMed handle = Entrez.esearch(db="pubmed", term="biopython", retmax=10) results = Entrez.read(handle) handle.close() print(f"Found {results['Count']} results") ``` ### 4. BLAST Operations (Bio.Blast) **Reference:** `references/blast.md` Use for: - Running BLAST searches via NCBI web services - Running local BLAST searches - Parsing BLAST XML output - Filtering results by E-value or identity - Extracting hit sequences **Quick example:** ```python from Bio.Blast import NCBIWWW, NCBIXML # Run BLAST search result_handle = NCBIWWW.qblast("blastn", "nt", "ATCGATCGATCG") blast_record = NCBIXML.read(result_handle) # Display top hits for alignment in blast_record.alignments[:5]: print(f"{alignment.title}: E-value={alignment.hsps[0].expect}") ``` ### 5. Structural Bioinformatics (Bio.PDB) **Reference:** `references/structure.md` Use for: - Parsing PDB and mmCIF structure files - Navigating protein structure hierarchy (SMCRA: Structure/Model/Chain/Residue/Atom) - Calculating distances, angles, and dihedrals - Secondary structure assignment (DSSP) - Structure superimposition and RMSD calculation - Extracting sequences from structures **Quick example:** ```python from Bio.PDB import PDBParser # Parse structure parser = PDBParser(QUIET=True) structure = parser.get_structure("1crn", "1crn.pdb") # Calculate distance between alpha carbons chain = structure[0]["A"] distance = chain[10]["CA"] - chain[20]["CA"] print(f"Distance: {distance:.2f} Å") ``` ### 6. Phylogenetics (Bio.Phylo) **Reference:** `references/phylogenetics.md` Use for: - Reading and writing phylogenetic trees (Newick, NEXUS, phyloXML) - Building trees from distance matrices or alignments - Tree manipulation (pruning, rerooting, ladderizing) - Calculating phylogenetic distances - Creating consensus trees - Visualizing trees **Quick example:** ```python from Bio import Phylo # Read and visualize tree tree = Phylo.read("tree.nwk", "newick") Phylo.draw_ascii(tree) # Calculate distance distance = tree.distance("Species_A", "Species_B") print(f"Distance: {distance:.3f}") ``` ### 7. Advanced Features **Reference:** `references/advanced.md` Use for: - **Sequence motifs** (Bio.motifs) - Finding and analyzing motif patterns - **Population genetics** (Bio.PopGen) - GenePop files, Fst calculations, Hardy-Weinberg tests - **Sequence utilities** (Bio.SeqUtils) - GC content, melting temperature, molecular weight, protein analysis - **Restriction analysis** (Bio.Restriction) - Finding restriction enzyme sites - **Clustering** (Bio.Cluster) - K-means and hierarchical clustering - **Genome diagrams** (GenomeDiagram) - Visualizing genomic features **Quick example:** ```python from Bio.SeqUtils import gc_fraction, molecular_weight from Bio.Seq import Seq seq = Seq("ATCGATCGATCG") print(f"GC content: {gc_fraction(seq):.2%}") print(f"Molecular weight: {molecular_weight(seq, seq_type='DNA'):.2f} g/mol") ``` ## General Workflow Guidelines ### Reading Documentation When a user asks about a specific Biopython task: 1. **Identify the relevant module** based on the task description 2. **Read the appropriate reference file** using the Read tool 3. **Extract relevant code patterns** and adapt them to the user's specific needs 4. **Combine multiple modules** when the task requires it Example search patterns for reference files: ```bash # Find information about specific functions rg -n "SeqIO.parse" references/sequence_io.md # Find examples of specific tasks rg -n "BLAST" references/blast.md # Find information about specific concepts rg -n "alignment" references/alignment.md ``` ### Writing Biopython Code Follow these principles when writing Biopython code: 1. **Import modules explicitly** ```python from Bio import SeqIO, Entrez from Bio.Seq import Seq ``` 2. **Set Entrez email** when using NCBI databases; load only `NCBI_API_KEY` from the environment if present ```python import os from Bio import Entrez Entrez.email = "your.email@example.com" Entrez.tool = "your_tool_name" if api_key := os.environ.get("NCBI_API_KEY"): Entrez.api_key = api_key ``` 3. **Use appropriate file formats** - Check which format best suits the task ```python # Common formats: "fasta", "genbank", "fastq", "clustal", "phylip" ``` 4. **Handle files properly** - Close handles after use or use context managers ```python with open("file.fasta") as handle: records = SeqIO.parse(handle, "fasta") ``` 5. **Use iterators for large files** - Avoid loading everything into memory ```python for record in SeqIO.parse("large_file.fasta", "fasta"): # Process one record at a time ``` 6. **Handle errors gracefully** - Network operations and file parsing can fail ```python from urllib.error import HTTPError try: handle = Entrez.efetch(db="nucleotide", id=accession) except HTTPError as e: print(f"Error: {e}") ``` ## Common Patterns ### Pattern 1: Fetch Sequence from GenBank ```python from Bio import Entrez, SeqIO Entrez.email = "your.email@example.com" # Fetch sequence handle = Entrez.efetch(db="nucleotide", id="EU490707", rettype="gb", retmode="text") record = SeqIO.read(handle, "genbank") handle.close() print(f"Description: {record.description}") print(f"Sequence length: {len(record.seq)}") ``` ### Pattern 2: Sequence Analysis Pipeline ```python from Bio import SeqIO from Bio.SeqUtils import gc_fraction for record in SeqIO.parse("sequences.fasta", "fasta"): # Calculate statistics gc = gc_fraction(record.seq) length = len(record.seq) # Find ORFs, translate, etc. protein = record.seq.translate() print(f"{record.id}: {length} bp, GC={gc:.2%}") ``` ### Pattern 3: BLAST and Fetch Top Hits ```python from Bio.Blast import NCBIWWW, NCBIXML from Bio import Entrez, SeqIO Entrez.email = "your.email@example.com" # Run BLAST result_handle = NCBIWWW.qblast("blastn", "nt", sequence) blast_record = NCBIXML.read(result_handle) # Get top hit accessions accessions = [aln.accession for aln in blast_record.alignments[:5]] # Fetch sequences for acc in accessions: handle = Entrez.efetch(db="nucleotide", id=acc, rettype="fasta", retmode="text") record = SeqIO.read(handle, "fasta") handle.close() print(f">{record.description}") ``` ### Pattern 4: Build Phylogenetic Tree from Sequences ```python from Bio import AlignIO, Phylo from Bio.Phylo.TreeConstruction import DistanceCalculator, DistanceTreeConstructor # Read alignment alignment = AlignIO.read("alignment.fasta", "fasta") # Calculate distances calculator = DistanceCalculator("identity") dm = calculator.get_distance(alignment) # Build tree constructor = DistanceTreeConstructor() tree = constructor.nj(dm) # Visualize Phylo.draw_ascii(tree) ``` ## Best Practices 1. **Always read relevant reference documentation** before writing code 2. **Use grep to search reference files** for specific functions or examples 3. **Validate file formats** before parsing 4. **Handle missing data gracefully** - Not all records have all fields 5. **Cache downloaded data** - Don't repeatedly download the same sequences 6. **Respect NCBI rate limits** - Use API keys, registered tool/email values for reusable software, and Entrez history/batching for large jobs 7. **Test with small datasets** before processing large files 8. **Keep Biopython updated** to get latest features and bug fixes 9. **Use appropriate genetic code tables** for translation 10. **Document analysis parameters** for reproducibility ## Troubleshooting Common Issues ### Issue: "No handlers could be found for logger 'Bio.Entrez'" **Solution:** This is just a warning. Set Entrez.email to suppress it. ### Issue: "HTTP Error 400" from NCBI **Solution:** Check that IDs/accessions are valid and properly formatted. ### Issue: "ValueError: EOF" when parsing files **Solution:** Verify file format matches the specified format string. ### Issue: Alignment fails with "sequences are not the same length" **Solution:** Ensure sequences are aligned before using AlignIO or MultipleSeqAlignment. ### Issue: BLAST searches are slow **Solution:** Use local BLAST for large-scale searches, or cache results. ### Issue: PDB parser warnings **Solution:** Use `PDBParser(QUIET=True)` to suppress warnings, or investigate structure quality. ### Issue: ImportError for Bio.HMM, Bio.MarkovModel, or Bio.Application **Solution:** These modules were removed in Biopython 1.86. Use [hmmlearn](https://pypi.org/project/hmmlearn/) for HMMs and the standard library `subprocess` module instead of `Bio.Application` CLI wrappers. ### Issue: PairwiseAligner returns fewer alignments after upgrading to 1.86+ **Solution:** The default gap score changed from 0 to -1 in 1.86, eliminating trivial tie alignments. Set `aligner.gap_score = 0` to restore the old behavior if needed (see `references/alignment.md`). ## Additional Resources - **Official Documentation**: https://biopython.org/docs/latest/ - **Tutorial**: https://biopython.org/docs/latest/Tutorial/ - **Cookbook**: https://biopython.org/docs/latest/Tutorial/ (advanced examples) - **GitHub**: https://github.com/biopython/biopython - **Release notes**: https://github.com/biopython/biopython/blob/master/NEWS.rst - **Deprecated APIs**: https://github.com/biopython/biopython/blob/master/DEPRECATED.rst - **Mailing List**: biopython@biopython.org ## Quick Reference To locate information in reference files, use these search patterns: ```bash # Search for specific functions rg -n "function_name" references/*.md # Find examples of specific tasks rg -n "example" references/sequence_io.md # Find all occurrences of a module rg -n "Bio.Seq" references/*.md ``` ## Summary Biopython provides comprehensive tools for computational molecular biology. When using this skill: 1. **Identify the task domain** (sequences, alignments, databases, BLAST, structures, phylogenetics, or advanced) 2. **Consult the appropriate reference file** in the `references/` directory 3. **Adapt code examples** to the specific use case 4. **Combine multiple modules** when needed for complex workflows 5. **Follow best practices** for file handling, error checking, and data management The modular reference documentation ensures detailed, searchable information for every major Biopython capability. ## Dónde encaja - Categoría: [Investigación](https://skillsagentes.com/categorias/investigacion.md) — Investigación estructurada, búsqueda de fuentes y síntesis. - Creador: [K-Dense-AI](https://skillsagentes.com/creators/k-dense-ai.md) — 163 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Citation Management](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/citation-management.md): Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas. - [Scientific Slides](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scientific-slides.md): Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual. - [Literature Review](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/literature-review.md): Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver). - [Infographics](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/infographics.md): Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo. - [Latex Posters](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/latex-posters.md): Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)