2 · DataFrames · lesson 8 of 20
groupBy & Aggregations
Compute sums, counts and stats across billions of rows.
Python
from pyspark.sql import functions as F
(df
.groupBy("country", "product")
.agg(
F.count("*").alias("orders"),
F.sum("amount").alias("revenue"),
F.avg("amount").alias("avg_ticket"),
F.countDistinct("customer_id").alias("customers"),
)
.orderBy(F.col("revenue").desc())
.show(20)
)WATCH OUT
groupBy always triggers a shuffle — data with the same key must land on the same executor. The 3D scene shows how rows are bucketed by hashed key.
Loading 3D scene…
Key takeaways
- ✓agg() takes multiple aggregations at once — always prefer it to chained groupBy calls.
- ✓countDistinct is expensive; approx_count_distinct is often good enough and 10-100× faster.
- ✓Every groupBy causes a shuffle across the network.