3 · Under the hood · lesson 15 of 20
Cache & Persist
Materialize a DataFrame once and reuse it across many actions.
Python
from pyspark import StorageLevel
base = (spark.read.parquet("/data/big")
.filter("year = 2024")
.repartition("country"))
base.cache() # MEMORY_AND_DISK by default
# or explicitly:
base.persist(StorageLevel.MEMORY_AND_DISK_SER)
base.count() # materializes and stores partitions
base.groupBy(...).show() # now fast — no re-scan
base.unpersist() # free memory when you're done- ▸MEMORY_ONLY — fastest, evicted under pressure.
- ▸MEMORY_AND_DISK — default, spills to disk when memory is tight.
- ▸*_SER variants — serialized, smaller footprint, slightly slower to read.
WATCH OUT
Caching is not free — it uses executor memory. Only cache DataFrames you'll reuse at least twice, and unpersist when you're done.
Loading 3D scene…
Key takeaways
- ✓cache() = persist(MEMORY_AND_DISK).
- ✓The DataFrame is only materialized on the next action.
- ✓Always unpersist() to release memory in long-running apps.