1 · Foundations · lesson 1 of 20
What is PySpark?
The Python API for Apache Spark — distributed compute for data that doesn't fit on one machine.
PySpark is the Python interface to Apache Spark, an open-source engine for large-scale data processing. You write Python code, but the actual work runs in parallel across a cluster of machines managed by Spark's JVM engine.
A Spark application has three roles: a Driver that runs your script and plans the work, a Cluster Manager that hands out resources, and Executors — worker processes on each node that run tasks and hold data partitions in memory.
The pipeline in one line
your_script.py → PySpark Driver → Cluster Manager → Executors → your dataTIP
Rotate the 3D scene below to see the Driver, Cluster Manager and Executors as separate stages that tasks flow through.
Loading 3D scene…
Key takeaways
- ✓PySpark = Python API on top of the Spark JVM engine (via Py4J).
- ✓Driver plans, Executors run, Cluster Manager coordinates.
- ✓Same code scales from a laptop to thousands of nodes.