# Scanpy > Pipeline estándar de análisis scRNA-seq: control de calidad, normalización, reducción de dimensionalidad (PCA/UMAP/t-SNE), clustering, expresión diferencial, visualización y conversión de formatos R a h5ad. Fuente: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scanpy Markdown: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scanpy.md Repositorio: https://github.com/K-Dense-AI/scientific-agent-skills Autor: K-Dense-AI Licencia: BSD-3-Clause Actualizado: el mes pasado Coste de contexto: 109 tok instalada, 3.9k tok al activarse, 30.4k tok con todos los archivos del bundle Bundle: 25 archivos, 119 KB Permisos que pide: ninguno declarado ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent claude-code # Cursor npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent cursor # Codex npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent codex # Gemini CLI npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent gemini # Windsurf npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent windsurf # Cline npx -y skills add K-Dense-AI/scientific-agent-skills --skill scanpy --agent cline ``` ## Qué hace - Ejecuta el pipeline completo de scRNA-seq: carga, QC, normalización, HVG, PCA, UMAP, clustering Leiden y marcadores - Incluye scripts CLI listos para cada paso (qc_analysis.py, preprocess.py, cluster.py, find_markers.py, etc.) - Convierte formatos de un solo comando desde 10X, CSV, loom o mtx a `.h5ad` - Convierte objetos R (Seurat, SingleCellExperiment `.rds`/`.RData`) a `.h5ad` para analizarlos en Scanpy - Genera visualizaciones publicables (UMAP, dotplot, heatmap) y agregados pseudobulk para expresión diferencial ## Cuándo usarla - Se analizan datos scRNA-seq en formato .h5ad, 10X o CSV - Se trabaja con datasets R-nativos (.rds, .RData, Seurat, SingleCellExperiment) que necesitan convertirse a .h5ad - Se identifican clusters de células, marcadores o se anotan tipos celulares - Se genera visualización publicable de un análisis de célula única ## Cuándo no - Se necesitan modelos de deep learning para single-cell (usar scvi-tools) - Son preguntas sobre el formato de datos AnnData (usar la skill anndata) ## Qué la activa - "Corre el pipeline completo de scRNA-seq sobre este archivo raw.h5ad" - "Convierte este objeto Seurat .rds a h5ad para analizarlo con Scanpy" - "Haz clustering Leiden a varias resoluciones y encuentra los marcadores por cluster" - "Genera un UMAP anotado por tipo celular de este dataset" ## Antes de instalar - Requiere Python 3.12+ y `scanpy[leiden]` (incluye python-igraph y leidenalg); convertir objetos R necesita herramientas de R aparte. - Necesita en el PATH: python ## Archivos - SKILL.md — 15 KB - assets/analysis_template.py — 10 KB - assets/celltype_mapping.json — 184 B - assets/gene_signatures.json — 355 B - assets/pipeline_config.json — 380 B - references/analysis_workflow.md — 7 KB - references/api_reference.md — 8 KB - references/plotting_guide.md — 10 KB - references/r_interop.md — 11 KB - references/standard_workflow.md — 6 KB - scripts/_common.py — 4 KB - scripts/annotate.py — 4 KB - scripts/batch_correct.py — 3 KB - scripts/cluster.py — 2 KB - scripts/convert.py — 1 KB - scripts/find_markers.py — 3 KB - scripts/inspect_data.py — 3 KB - scripts/plot.py — 4 KB - scripts/preprocess.py — 4 KB - scripts/pseudobulk.py — 3 KB - scripts/qc_analysis.py — 5 KB - scripts/reduce_dimensions.py — 3 KB - scripts/run_pipeline.py — 8 KB - scripts/score_genes.py — 4 KB - scripts/subset.py — 2 KB ## SKILL.md Reproducido tal cual desde K-Dense-AI/scientific-agent-skills bajo BSD-3-Clause. Esta sección es el documento original y está en inglés. # Scanpy: Single-Cell Analysis ## Overview Scanpy is a scalable Python toolkit for analyzing single-cell RNA-seq data, built on AnnData. Apply this skill for complete single-cell workflows including quality control, normalization, dimensionality reduction, clustering, marker gene identification, visualization, and trajectory analysis. Current stable release: **scanpy 1.12.x** (January 2026). ## Installation Requires Python **3.12+** (scanpy 1.12 dropped Python ≤3.11) and anndata **≥0.10**. ```bash uv pip install "scanpy[leiden]" ``` The `[leiden]` extra installs `python-igraph` and `leidenalg`, required for Leiden clustering. For reproducible environments, pin a version: `uv pip install "scanpy[leiden]==1.12.1"`. For large or out-of-core datasets, many functions support [Dask](https://docs.dask.org/) arrays (experimental): ```bash uv pip install "scanpy[leiden]" dask ``` See the [Using dask with Scanpy](https://scanpy.scverse.org/en/stable/tutorials/experimental/dask.html) tutorial. For GPU-accelerated scanpy-like operations, use [rapids-singlecell](https://rapids-singlecell.readthedocs.io/) as a separate package. If the input is an R-native single-cell object (`.rds`, `.RData`, Seurat, or SingleCellExperiment), first convert it to `.h5ad` with R tooling, then load it with Scanpy. Read `references/r_interop.md` for agent-run installation and conversion instructions across macOS, Linux, and Windows. For AnnData structure and I/O details, use the **anndata** skill. For probabilistic models and batch correction, use **scvi-tools**. ## When to Use This Skill This skill should be used when: - Analyzing single-cell RNA-seq data (.h5ad, 10X, CSV formats) - Working with R-friendly single-cell datasets (`.rds`, `.RData`, Seurat, SingleCellExperiment) that need conversion to `.h5ad` - Performing quality control on scRNA-seq datasets - Creating UMAP, t-SNE, or PCA visualizations - Identifying cell clusters and finding marker genes - Annotating cell types based on gene expression - Conducting trajectory inference or pseudotime analysis - Generating publication-quality single-cell plots ## Script Toolkit (prefer these over writing code from scratch) This skill bundles ready-to-run CLI scripts in `scripts/` for every common step. **Run these instead of hand-writing scanpy code** — they handle file loading by extension, figure setup, sensible defaults, raw-count preservation, and progress logging. Each reads and writes `.h5ad`, so they chain together, and each has its own `--help`. Only drop down to writing scanpy code when a task isn't covered by a script or needs unusual customization. All scripts use a shared `scripts/_common.py` helper (loading, saving, figure config) — keep it alongside the others. Run from the skill directory or pass full paths; figures default to `./figures/`. | Script | Purpose | Typical call | |--------|---------|--------------| | `run_pipeline.py` | **Full workflow in one command**: load → QC → normalize → HVG → PCA → (batch) → UMAP → Leiden → markers | `python scripts/run_pipeline.py raw.h5ad -o processed.h5ad` | | `inspect_data.py` | Summarize an unknown dataset (shape, obs/var, layers, what's already computed, raw vs normalized) | `python scripts/inspect_data.py data.h5ad` | | `convert.py` | Load any format (10x dir/.h5, csv, loom, mtx) and write `.h5ad` | `python scripts/convert.py 10x_dir/ -o data.h5ad` | | `qc_analysis.py` | QC metrics, before/after plots, filtering, optional Scrublet doublets | `python scripts/qc_analysis.py raw.h5ad -o qc.h5ad --scrublet` | | `preprocess.py` | Normalize, log1p, HVG, optional scale/regress (keeps `counts` layer + `raw`) | `python scripts/preprocess.py qc.h5ad -o norm.h5ad` | | `reduce_dimensions.py` | PCA + variance plot, neighbors, UMAP, optional t-SNE | `python scripts/reduce_dimensions.py norm.h5ad -o red.h5ad` | | `batch_correct.py` | Integration: harmony / bbknn / combat | `python scripts/batch_correct.py red.h5ad -o int.h5ad --method harmony --batch-key sample` | | `cluster.py` | Leiden (or louvain) at one or many resolutions | `python scripts/cluster.py red.h5ad -o clu.h5ad --resolution 0.3 0.6 1.0` | | `find_markers.py` | `rank_genes_groups` + per-group CSVs + marker plots | `python scripts/find_markers.py clu.h5ad --groupby leiden -o clu.h5ad` | | `annotate.py` | Map clusters → cell types from JSON/CSV; optional marker reference dotplot | `python scripts/annotate.py clu.h5ad -o ann.h5ad --mapping map.json` | | `score_genes.py` | Score gene signatures (JSON) and/or cell-cycle phase | `python scripts/score_genes.py ann.h5ad -o scored.h5ad --gene-sets sigs.json` | | `pseudobulk.py` | Aggregate counts by sample × cell type → matrix for pydeseq2 | `python scripts/pseudobulk.py ann.h5ad --by sample cell_type --out-prefix pb` | | `subset.py` | Subset by obs values or gene list (optionally clear stale embeddings) | `python scripts/subset.py ann.h5ad -o tcells.h5ad --obs cell_type --keep "T cells"` | | `plot.py` | Generate umap/tsne/pca/violin/dotplot/heatmap/etc. from a processed object | `python scripts/plot.py ann.h5ad --kind dotplot --genes CD3D CD14 --groupby cell_type` | ### One-shot end-to-end run ```bash # Counts → clustered, marker-annotated object + figures + marker CSVs python scripts/run_pipeline.py raw.h5ad -o processed.h5ad \ --resolution 0.5 --n-top-genes 2000 --scrublet # With multi-sample integration: python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --batch-key sample --batch-method harmony # Reproducible parameters via JSON (keys mirror flag names with underscores): python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --config params.json ``` ### Step-by-step chain (when you need to inspect/iterate between stages) ```bash python scripts/qc_analysis.py raw.h5ad -o qc.h5ad --scrublet python scripts/preprocess.py qc.h5ad -o norm.h5ad --n-top-genes 2000 python scripts/reduce_dimensions.py norm.h5ad -o red.h5ad --n-pcs 40 python scripts/cluster.py red.h5ad -o clu.h5ad --resolution 0.3 0.5 0.8 python scripts/find_markers.py clu.h5ad -o clu.h5ad --groupby leiden --use-raw # inspect results/markers/*.csv, decide labels, write a mapping JSON, then: python scripts/annotate.py clu.h5ad -o ann.h5ad --mapping celltypes.json ``` The sections below document the underlying scanpy calls each script performs — read them when customizing beyond the script flags. ## Quick Start ### Basic Import and Setup ```python import scanpy as sc import pandas as pd import numpy as np # Configure settings sc.settings.verbosity = 3 sc.settings.set_figure_params(dpi=80, facecolor='white') sc.settings.figdir = './figures/' sc.settings.autosave = True # Preferred over per-plot save= (deprecated in scanpy 1.12) ``` ### Loading Data ```python # From 10X Genomics adata = sc.read_10x_mtx('path/to/data/') adata = sc.read_10x_h5('path/to/data.h5') # From h5ad (AnnData format) adata = sc.read_h5ad('path/to/data.h5ad') # From CSV adata = sc.read_csv('path/to/data.csv') ``` For R-native files, do not try to parse Seurat `.rds` directly in Python. Convert first: ```bash # See references/r_interop.md for installing R and conversion packages. Rscript convert_rds_to_h5ad.R input.rds output.h5ad ``` ```python adata = sc.read_h5ad('output.h5ad') ``` ### Understanding AnnData Structure The AnnData object is the core data structure in scanpy: ```python adata.X # Expression matrix (cells × genes) adata.obs # Cell metadata (DataFrame) adata.var # Gene metadata (DataFrame) adata.uns # Unstructured annotations (dict) adata.obsm # Multi-dimensional cell data (PCA, UMAP) adata.raw # Raw data backup # Access cell and gene names adata.obs_names # Cell barcodes adata.var_names # Gene names ``` ## Standard Analysis Workflow The seven steps, with code and the parameters that matter at each, are in [references/analysis_workflow.md](references/analysis_workflow.md): 1. **Quality control** — filter cells and genes; inspect mitochondrial fraction and counts before choosing thresholds rather than copying defaults. 2. **Normalization and preprocessing** — normalize, log-transform, select highly variable genes, and keep `.raw` for later plotting. 3. **Dimensionality reduction** — PCA, then the neighbour graph, then UMAP. 4. **Clustering** — Leiden at a resolution chosen for the question, not the default. 5. **Marker gene identification** — ranked genes per cluster. 6. **Cell type annotation** — mapping clusters to types from markers. 7. **Save results** — writing the annotated `AnnData`. Common follow-on tasks — publication plots, trajectory inference, pseudobulk differential expression between conditions, gene set scoring, and batch correction — are in the same file. See also [references/standard_workflow.md](references/standard_workflow.md) and [references/plotting_guide.md](references/plotting_guide.md). ## Key Parameters to Adjust ### Quality Control - `min_genes`: Minimum genes per cell (typically 200-500) - `min_cells`: Minimum cells per gene (typically 3-10) - `pct_counts_mt`: Mitochondrial threshold (typically 5-20%) ### Normalization - `target_sum`: Target counts per cell (default 1e4) ### Feature Selection - `n_top_genes`: Number of HVGs (typically 2000-3000) - `min_mean`, `max_mean`, `min_disp`: HVG selection parameters ### Dimensionality Reduction - `n_pcs`: Number of principal components (check variance ratio plot) - `n_neighbors`: Number of neighbors (typically 10-30) ### Clustering - `resolution`: Clustering granularity (0.4-1.2, higher = more clusters) ## Common Pitfalls and Best Practices 1. **Always save raw counts**: `adata.raw = adata` before filtering genes 2. **Check QC plots carefully**: Adjust thresholds based on dataset quality 3. **Use Leiden clustering**: `sc.tl.louvain` is deprecated in scanpy 1.12 4. **Try multiple clustering resolutions**: Find optimal granularity 5. **Validate cell type annotations**: Use multiple marker genes 6. **Use `use_raw=True` for gene expression plots**: Shows normalized counts from `.raw` 7. **Check PCA variance ratio**: Determine optimal number of PCs 8. **Save intermediate results**: Long workflows can fail partway through 9. **Pseudobulk for DE**: Do not treat `rank_genes_groups` p-values as rigorous DE between conditions 10. **Save plots via settings**: Use `sc.settings.autosave` instead of deprecated `save=` on plot functions 11. **Convert R objects before Scanpy**: Use R packages to convert Seurat or SingleCellExperiment `.rds` files to `.h5ad`, preserving counts, metadata, and gene identifiers ## Bundled Resources ### scripts/ (CLI toolkit) A composable set of `.h5ad`-in/`.h5ad`-out scripts covering the whole workflow plus a one-command end-to-end pipeline. See the **Script Toolkit** section above for the full table and chaining examples. Each script has `--help`. Files: - `_common.py` — shared loading/saving/figure helpers imported by the others (not a CLI) - `run_pipeline.py` — full pipeline in one command (flags or `--config` JSON) - `inspect_data.py`, `convert.py` — explore and load/convert any input format - `qc_analysis.py`, `preprocess.py`, `reduce_dimensions.py`, `batch_correct.py`, `cluster.py` — pipeline steps - `find_markers.py`, `annotate.py`, `score_genes.py`, `pseudobulk.py` — markers, annotation, scoring, DE prep - `subset.py`, `plot.py` — subset by metadata/genes; generate any standard plot **Default to these scripts before writing scanpy code from scratch.** ### references/standard_workflow.md Complete step-by-step workflow with detailed explanations and code examples for: - Data loading and setup - Quality control with visualization - Normalization and scaling - Feature selection - Dimensionality reduction (PCA, UMAP, t-SNE) - Clustering (Leiden) - Doublet detection (scrublet) and pseudobulk aggregation - Marker gene identification - Cell type annotation - Trajectory inference - Differential expression Read this reference when performing a complete analysis from scratch. ### references/api_reference.md Quick reference guide for scanpy functions organized by module: - Reading/writing data (`sc.read_*`, `adata.write_*`) - Preprocessing (`sc.pp.*`) - Tools (`sc.tl.*`) - Plotting (`sc.pl.*`) - AnnData structure and manipulation - Settings and utilities Use this for quick lookup of function signatures and common parameters. ### references/plotting_guide.md Comprehensive visualization guide including: - Quality control plots - Dimensionality reduction visualizations - Clustering visualizations - Marker gene plots (heatmaps, dot plots, violin plots) - Trajectory and pseudotime plots - Publication-quality customization - Multi-panel figures - Color palettes and styling Consult this when creating publication-ready figures. ### references/r_interop.md Agent runbook for installing R on macOS, Linux, and Windows, installing CRAN/Bioconductor conversion packages, inspecting `.rds`/`.RData` inputs, converting Seurat or SingleCellExperiment objects to `.h5ad`, and validating the result in Scanpy. ### assets/analysis_template.py Complete analysis template providing a full workflow from data loading through cell type annotation. Copy and customize this template for new analyses: ```bash cp assets/analysis_template.py my_analysis.py # Edit parameters and run python my_analysis.py ``` The template includes all standard steps with configurable parameters and helpful comments. ### assets/ JSON templates Edit-and-pass templates so you don't author config/mappings from scratch: - `assets/pipeline_config.json` — parameter set for `run_pipeline.py --config` - `assets/celltype_mapping.json` — cluster → cell-type map for `annotate.py --mapping` - `assets/gene_signatures.json` — gene-set signatures for `score_genes.py --gene-sets` ## Additional Resources - **Official scanpy documentation**: https://scanpy.scverse.org/en/stable/ - **Scanpy tutorials**: https://scanpy.scverse.org/en/stable/tutorials/index.html - **Release notes**: https://scanpy.scverse.org/en/stable/release-notes/index.html - **scverse ecosystem**: https://scverse.org/ (related tools: squidpy, scvi-tools, cellrank) - **R interoperability**: https://www.bioconductor.org/packages/release/bioc/html/zellkonverter.html and https://mojaveazure.github.io/seurat-disk/ - **Best practices**: Luecken & Theis (2019) "Current best practices in single-cell RNA-seq" ## Tips for Effective Analysis 1. **Start with the template**: Use `assets/analysis_template.py` as a starting point 2. **Run QC script first**: Use `scripts/qc_analysis.py` for initial filtering 3. **Consult references as needed**: Load workflow and API references into context 4. **Iterate on clustering**: Try multiple resolutions and visualization methods 5. **Validate biologically**: Check marker genes match expected cell types 6. **Document parameters**: Record QC thresholds and analysis settings 7. **Save checkpoints**: Write intermediate results at key steps ## Dónde encaja - Categoría: [Datos y analítica](https://skillsagentes.com/categorias/datos-analitica.md) — Consulta, limpia y visualiza datos sin salir del agente. - Creador: [K-Dense-AI](https://skillsagentes.com/creators/k-dense-ai.md) — 163 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Citation Management](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/citation-management.md): Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas. - [Scientific Slides](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scientific-slides.md): Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual. - [Literature Review](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/literature-review.md): Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver). - [Infographics](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/infographics.md): Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo. - [Latex Posters](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/latex-posters.md): Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)