4 · Advanced · lesson 19 of 20

Running on a Cluster

From spark-submit flags to why your job is slow.

shell
spark-submit \
    --master yarn \
    --deploy-mode cluster \
    --num-executors 20 \
    --executor-cores 4 \
    --executor-memory 8g \
    --conf spark.sql.shuffle.partitions=400 \
    --conf spark.sql.adaptive.enabled=true \
    my_job.py
  • ▸num-executors × executor-cores = total parallel tasks.
  • ▸executor-memory is split between execution, storage (cache) and overhead.
  • ▸spark.sql.adaptive.enabled=true lets Spark re-optimize at runtime — turn it on.
  • ▸Cluster managers: YARN, Kubernetes, Standalone, Mesos (deprecated).
NOTE
The 3D scene shows hundreds of tasks orbiting a few carrier executors — Spark's scheduler multiplexes many tasks onto a small pool of cores.
Loading 3D scene…
Key takeaways
  • ✓Right-size executors: too big = wasted cores, too small = overhead.
  • ✓Adaptive Query Execution (AQE) auto-tunes shuffle partitions and join strategies.
  • ✓Always start from the Spark UI when a job is slow — don't guess.