4 · Advanced · lesson 17 of 20
UDFs & pandas UDFs
When built-in functions aren't enough — and how to keep it fast.
A Python UDF ships each row to the Python interpreter — correct, but 10-100× slower than a built-in function. A pandas UDF ships whole Arrow batches to a vectorized pandas function, closing most of the gap.
Python
from pyspark.sql import functions as F
from pyspark.sql.types import DoubleType
import pandas as pd
# Regular Python UDF — slow
@F.udf(DoubleType())
def celsius(f):
return (f - 32) * 5.0 / 9.0
# pandas UDF — vectorized, uses Apache Arrow
@F.pandas_udf(DoubleType())
def celsius_fast(f: pd.Series) -> pd.Series:
return (f - 32) * 5.0 / 9.0
df.withColumn("c", celsius_fast("temp_f"))WATCH OUT
Always check for a built-in first. F.expr, F.when, F.regexp_extract and friends cover 90% of what people write UDFs for.
Loading 3D scene…
Key takeaways
- ✓Prefer built-in functions → pandas UDF → plain Python UDF, in that order.
- ✓pandas UDFs need PyArrow installed on every executor.
- ✓UDFs are opaque to Catalyst — it can't push filters through them.