3 · Under the hood · lesson 14 of 20
Catalyst & Tungsten
Why DataFrame code often beats hand-written RDD code.
Catalyst is Spark's query optimizer. It rewrites your DataFrame plan: pushing filters into the file scan, pruning unused columns, reordering joins, combining projections. Tungsten is the execution layer that turns the optimized plan into cache-friendly, whole-stage-code-generated bytecode.
Python
df.explain(mode="formatted")
# == Physical Plan ==
# * HashAggregate(keys=[country], functions=[sum(amount)])
# +- Exchange hashpartitioning(country, 200)
# +- * HashAggregate(keys=[country], functions=[partial_sum(amount)])
# +- * Project [country, amount]
# +- * Filter (amount > 100)
# +- FileScan parquet [country, amount] PushedFilters: [GreaterThan(amount, 100)]TIP
The '*' in the physical plan means whole-stage code generation is active — Spark fused those operators into a single tight loop.
Key takeaways
- ✓Catalyst optimizes; Tungsten executes on efficient binary memory.
- ✓df.explain() is the fastest way to see if predicates and columns are being pushed down.
- ✓Prefer the DataFrame API so you get these optimizations for free.