4 · Advanced · lesson 20 of 20

Performance Playbook

The checklist to run through before every tuning session.

  • ▸Read only the columns you need — Parquet + select() enables column pruning.
  • ▸Filter as early as possible — Catalyst pushes predicates into the scan.
  • ▸Broadcast small dimensions instead of shuffling them.
  • ▸Watch out for data skew — one huge key can stall a stage. Salt the key or use skew hints.
  • ▸Cache reused intermediates; unpersist when done.
  • ▸Prefer built-in functions over UDFs.
  • ▸Turn on AQE (spark.sql.adaptive.enabled=true) — free wins.
  • ▸Right-size spark.sql.shuffle.partitions for your data volume.
TIP
The single fastest debugging move: open the Spark UI, sort stages by duration, and click the slowest one. 90% of tuning starts there.
Loading 3D scene…
Key takeaways
  • ✓Most PySpark performance work is about reducing shuffle and skew.
  • ✓The Spark UI answers 'where is time going?' faster than any code review.
  • ✓Adaptive Query Execution is free — enable it everywhere.