2 · DataFrames · lesson 9 of 20

Joins

Inner, left, right, full, semi, anti — plus the join Spark does for free.

Python
orders.join(customers, on="customer_id", how="inner")

orders.join(
    customers,
    orders.customer_id == customers.id,
    how="left",
)

# Broadcast a small side to skip the shuffle
from pyspark.sql.functions import broadcast
orders.join(broadcast(countries), "country_code")
  • ▸how: inner, left, right, full, left_semi, left_anti, cross.
  • ▸Semi = rows in left that have a match; anti = rows in left with NO match.
  • ▸broadcast() ships the small DataFrame to every executor — huge speedup when one side is < ~10 MB.
Key takeaways
  • ✓Standard joins shuffle both sides on the join key.
  • ✓Broadcast joins skip the shuffle when one side is small.
  • ✓Semi/anti joins are the clean way to say 'filter by existence in another table'.