1 · Foundations · lesson 3 of 20
RDD vs DataFrame vs Dataset
Three APIs, one engine — and why DataFrames are almost always the right choice.
Spark started with RDDs (Resilient Distributed Datasets): immutable collections of Python objects split across the cluster. RDDs give total control but no schema and no automatic optimization.
DataFrames added a schema on top — rows and named, typed columns, like a distributed table. Spark's Catalyst optimizer can then rewrite your query for you.
Python
# RDD — low level, arbitrary Python objects
rdd = spark.sparkContext.parallelize([("Ada", 36), ("Grace", 85)])
rdd.map(lambda t: (t[0].upper(), t[1] + 1)).collect()
# DataFrame — schema, SQL, Catalyst optimizer
df = spark.createDataFrame(
[("Ada", 36), ("Grace", 85)],
schema=["name", "age"],
)
df.selectExpr("upper(name) AS name", "age + 1 AS age").show()TIP
Rule of thumb: reach for DataFrames first. Drop to RDDs only when you truly need per-row Python control that the DataFrame API can't express.
Loading 3D scene…
Key takeaways
- ✓RDD = distributed collection of Python objects, no schema.
- ✓DataFrame = distributed table with a schema; optimized by Catalyst.
- ✓Dataset is JVM-only (Scala/Java) — Python users get DataFrames.