# Spark Optimization > Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines. Source: https://skillsagentes.com/skills/wshobson/agents/spark-optimization Repository: https://github.com/wshobson/agents Author: wshobson License: MIT Updated: hace 2 meses Context cost: 48 tok installed, 790 tok once triggered, 3.3k tok with every bundled file Bundle: 2 files, 13 KB Permissions requested: none declared ## Install ```bash npx -y skills add wshobson/agents --skill spark-optimization --agent claude-code ``` ## What it does - Configura Spark con AQE, coalescing de particiones y skew join habilitados - Aplica Kryo serializer y formatos columnar (Parquet/Delta) en lecturas y escrituras - Ajusta particionamiento, memoria de ejecutores y broadcast joins para tablas pequeñas - Detecta y corrige data skew, spills y presión de GC vía Spark UI ## Use it when - Optimizar jobs de Spark lentos - Ajustar configuración de memoria y ejecutores - Implementar estrategias de particionamiento eficientes - Depurar problemas de rendimiento o escalar pipelines de Spark ## What triggers it - "Mi job de Spark está muy lento, ayúdame a optimizarlo" - "¿Cómo configuro el particionamiento y la memoria de executors en Spark?" - "Tengo data skew en un join de Spark, ¿cómo lo soluciono?" ## Before you install - Requiere un entorno con Apache Spark y PySpark configurado. ## Files - SKILL.md — 3 KB - references/details.md — 10 KB ## SKILL.md Reproduced verbatim from wshobson/agents under MIT. This section is the upstream document and is in English. # Apache Spark Optimization Production patterns for optimizing Apache Spark jobs including partitioning strategies, memory management, shuffle optimization, and performance tuning. ## When to Use This Skill - Optimizing slow Spark jobs - Tuning memory and executor configuration - Implementing efficient partitioning strategies - Debugging Spark performance issues - Scaling Spark pipelines for large datasets - Reducing shuffle and data skew ## Core Concepts ### 1. Spark Execution Model ``` Driver Program ↓ Job (triggered by action) ↓ Stages (separated by shuffles) ↓ Tasks (one per partition) ``` ### 2. Key Performance Factors | Factor | Impact | Solution | | ----------------- | --------------------- | ----------------------------- | | **Shuffle** | Network I/O, disk I/O | Minimize wide transformations | | **Data Skew** | Uneven task duration | Salting, broadcast joins | | **Serialization** | CPU overhead | Use Kryo, columnar formats | | **Memory** | GC pressure, spills | Tune executor memory | | **Partitions** | Parallelism | Right-size partitions | ## Quick Start ```python from pyspark.sql import SparkSession from pyspark.sql import functions as F # Create optimized Spark session spark = (SparkSession.builder .appName("OptimizedJob") .config("spark.sql.adaptive.enabled", "true") .config("spark.sql.adaptive.coalescePartitions.enabled", "true") .config("spark.sql.adaptive.skewJoin.enabled", "true") .config("spark.serializer", "org.apache.spark.serializer.KryoSerializer") .config("spark.sql.shuffle.partitions", "200") .getOrCreate()) # Read with optimized settings df = (spark.read .format("parquet") .option("mergeSchema", "false") .load("s3://bucket/data/")) # Efficient transformations result = (df .filter(F.col("date") >= "2024-01-01") .select("id", "amount", "category") .groupBy("category") .agg(F.sum("amount").alias("total"))) result.write.mode("overwrite").parquet("s3://bucket/output/") ``` ## Detailed patterns and worked examples Detailed pattern documentation lives in `references/details.md`. Read that file when the navigation tier above is insufficient. ## Best Practices ### Do's - **Enable AQE** - Adaptive query execution handles many issues - **Use Parquet/Delta** - Columnar formats with compression - **Broadcast small tables** - Avoid shuffle for small joins - **Monitor Spark UI** - Check for skew, spills, GC - **Right-size partitions** - 128MB - 256MB per partition ### Don'ts - **Don't collect large data** - Keep data distributed - **Don't use UDFs unnecessarily** - Use built-in functions - **Don't over-cache** - Memory is limited - **Don't ignore data skew** - It dominates job time - **Don't use `.count()` for existence** - Use `.take(1)` or `.isEmpty()` --- Skills Agentes — https://skillsagentes.com/skills/wshobson/agents/spark-optimization