4 · Advanced · lesson 17 of 20

UDFs & pandas UDFs

When built-in functions aren't enough — and how to keep it fast.

A Python UDF ships each row to the Python interpreter — correct, but 10-100× slower than a built-in function. A pandas UDF ships whole Arrow batches to a vectorized pandas function, closing most of the gap.

Python
from pyspark.sql import functions as F
from pyspark.sql.types import DoubleType
import pandas as pd

# Regular Python UDF — slow
@F.udf(DoubleType())
def celsius(f):
    return (f - 32) * 5.0 / 9.0

# pandas UDF — vectorized, uses Apache Arrow
@F.pandas_udf(DoubleType())
def celsius_fast(f: pd.Series) -> pd.Series:
    return (f - 32) * 5.0 / 9.0

df.withColumn("c", celsius_fast("temp_f"))
WATCH OUT
Always check for a built-in first. F.expr, F.when, F.regexp_extract and friends cover 90% of what people write UDFs for.
Loading 3D scene…
Key takeaways
  • ✓Prefer built-in functions → pandas UDF → plain Python UDF, in that order.
  • ✓pandas UDFs need PyArrow installed on every executor.
  • ✓UDFs are opaque to Catalyst — it can't push filters through them.