1 · Foundations · lesson 3 of 20

RDD vs DataFrame vs Dataset

Three APIs, one engine — and why DataFrames are almost always the right choice.

Spark started with RDDs (Resilient Distributed Datasets): immutable collections of Python objects split across the cluster. RDDs give total control but no schema and no automatic optimization.

DataFrames added a schema on top — rows and named, typed columns, like a distributed table. Spark's Catalyst optimizer can then rewrite your query for you.

Python
# RDD — low level, arbitrary Python objects
rdd = spark.sparkContext.parallelize([("Ada", 36), ("Grace", 85)])
rdd.map(lambda t: (t[0].upper(), t[1] + 1)).collect()

# DataFrame — schema, SQL, Catalyst optimizer
df = spark.createDataFrame(
    [("Ada", 36), ("Grace", 85)],
    schema=["name", "age"],
)
df.selectExpr("upper(name) AS name", "age + 1 AS age").show()
TIP
Rule of thumb: reach for DataFrames first. Drop to RDDs only when you truly need per-row Python control that the DataFrame API can't express.
Loading 3D scene…
Key takeaways
  • ✓RDD = distributed collection of Python objects, no schema.
  • ✓DataFrame = distributed table with a schema; optimized by Catalyst.
  • ✓Dataset is JVM-only (Scala/Java) — Python users get DataFrames.