2 · DataFrames · lesson 8 of 20

groupBy & Aggregations

Compute sums, counts and stats across billions of rows.

Python
from pyspark.sql import functions as F

(df
    .groupBy("country", "product")
    .agg(
        F.count("*").alias("orders"),
        F.sum("amount").alias("revenue"),
        F.avg("amount").alias("avg_ticket"),
        F.countDistinct("customer_id").alias("customers"),
    )
    .orderBy(F.col("revenue").desc())
    .show(20)
)
WATCH OUT
groupBy always triggers a shuffle — data with the same key must land on the same executor. The 3D scene shows how rows are bucketed by hashed key.
Loading 3D scene…
Key takeaways
  • ✓agg() takes multiple aggregations at once — always prefer it to chained groupBy calls.
  • ✓countDistinct is expensive; approx_count_distinct is often good enough and 10-100× faster.
  • ✓Every groupBy causes a shuffle across the network.