2 · DataFrames · lesson 9 of 20
Joins
Inner, left, right, full, semi, anti — plus the join Spark does for free.
Python
orders.join(customers, on="customer_id", how="inner")
orders.join(
customers,
orders.customer_id == customers.id,
how="left",
)
# Broadcast a small side to skip the shuffle
from pyspark.sql.functions import broadcast
orders.join(broadcast(countries), "country_code")- ▸how: inner, left, right, full, left_semi, left_anti, cross.
- ▸Semi = rows in left that have a match; anti = rows in left with NO match.
- ▸broadcast() ships the small DataFrame to every executor — huge speedup when one side is < ~10 MB.
Key takeaways
- ✓Standard joins shuffle both sides on the join key.
- ✓Broadcast joins skip the shuffle when one side is small.
- ✓Semi/anti joins are the clean way to say 'filter by existence in another table'.