4 · Advanced · lesson 20 of 20
Performance Playbook
The checklist to run through before every tuning session.
- ▸Read only the columns you need — Parquet + select() enables column pruning.
- ▸Filter as early as possible — Catalyst pushes predicates into the scan.
- ▸Broadcast small dimensions instead of shuffling them.
- ▸Watch out for data skew — one huge key can stall a stage. Salt the key or use skew hints.
- ▸Cache reused intermediates; unpersist when done.
- ▸Prefer built-in functions over UDFs.
- ▸Turn on AQE (spark.sql.adaptive.enabled=true) — free wins.
- ▸Right-size spark.sql.shuffle.partitions for your data volume.
TIP
The single fastest debugging move: open the Spark UI, sort stages by duration, and click the slowest one. 90% of tuning starts there.
Loading 3D scene…
Key takeaways
- ✓Most PySpark performance work is about reducing shuffle and skew.
- ✓The Spark UI answers 'where is time going?' faster than any code review.
- ✓Adaptive Query Execution is free — enable it everywhere.