3 · Under the hood · lesson 14 of 20

Catalyst & Tungsten

Why DataFrame code often beats hand-written RDD code.

Catalyst is Spark's query optimizer. It rewrites your DataFrame plan: pushing filters into the file scan, pruning unused columns, reordering joins, combining projections. Tungsten is the execution layer that turns the optimized plan into cache-friendly, whole-stage-code-generated bytecode.

Python
df.explain(mode="formatted")

# == Physical Plan ==
# * HashAggregate(keys=[country], functions=[sum(amount)])
#   +- Exchange hashpartitioning(country, 200)
#      +- * HashAggregate(keys=[country], functions=[partial_sum(amount)])
#         +- * Project [country, amount]
#            +- * Filter (amount > 100)
#               +- FileScan parquet [country, amount] PushedFilters: [GreaterThan(amount, 100)]
TIP
The '*' in the physical plan means whole-stage code generation is active — Spark fused those operators into a single tight loop.
Key takeaways
  • ✓Catalyst optimizes; Tungsten executes on efficient binary memory.
  • ✓df.explain() is the fastest way to see if predicates and columns are being pushed down.
  • ✓Prefer the DataFrame API so you get these optimizations for free.