3 · Under the hood · lesson 15 of 20

Cache & Persist

Materialize a DataFrame once and reuse it across many actions.

Python
from pyspark import StorageLevel

base = (spark.read.parquet("/data/big")
                 .filter("year = 2024")
                 .repartition("country"))

base.cache()                                   # MEMORY_AND_DISK by default
# or explicitly:
base.persist(StorageLevel.MEMORY_AND_DISK_SER)

base.count()          # materializes and stores partitions
base.groupBy(...).show()   # now fast — no re-scan

base.unpersist()      # free memory when you're done
  • ▸MEMORY_ONLY — fastest, evicted under pressure.
  • ▸MEMORY_AND_DISK — default, spills to disk when memory is tight.
  • ▸*_SER variants — serialized, smaller footprint, slightly slower to read.
WATCH OUT
Caching is not free — it uses executor memory. Only cache DataFrames you'll reuse at least twice, and unpersist when you're done.
Loading 3D scene…
Key takeaways
  • ✓cache() = persist(MEMORY_AND_DISK).
  • ✓The DataFrame is only materialized on the next action.
  • ✓Always unpersist() to release memory in long-running apps.