Pandas UDFs (Vectorized UDFs)

A pandas 16,086 UDF (@pandas_udf with type hints) receives a whole Arrow 129 batch as a pandas Series and returns one, so it runs once per batch with NumPy doing the work; Spark 4.1 129 's Arrow UDFs (@arrow_udf) hand you the pyarrow.Array itself. The benchmark applies one rule five ways to four million amounts held in executor storage.

One rule, five implementations, timedJavaScript
import timeit
import pandas as pd
import pyarrow as pa
import pyarrow.compute as pc
from pyspark.sql.functions import arrow_udf, pandas_udf, udf
def net(total):                                     # one rule, five implementations
    return None if total is None else round(total / 1.06, 2)
@pandas_udf("double")
def net_pandas(total: pd.Series) -> pd.Series:      # a whole Arrow batch as a Series
    return (total / 1.06).round(2)
@arrow_udf("double")
def net_arrow(total: pa.Array) -> pa.Array:         # the Arrow batch itself (Spark 4.1+)
    return pc.round(pc.divide(total, 1.06), 2)
amounts = (orders.crossJoin(spark.range(4))           # 4 million amounts
           .select(F.col("total").cast("double").alias("total"))
           .localCheckpoint())                      # materialized: time the UDF, not the scan
variants = {"built-in F.round": F.round(F.col("total") / 1.06, 2),
            "Python UDF, pickled": udf(net, "double", useArrow=False)("total"),
            "Python UDF, Arrow": udf(net, "double", useArrow=True)("total"),
            "pandas UDF": net_pandas("total"),
            "Arrow UDF": net_arrow("total")}
base = None
for name, col in variants.items():
    q = amounts.select(F.sum(col).alias("s"))
    secs = min(timeit.timeit(q.first, number=1) for _ in range(3))     # best of three
    base = base or secs                                                # built-in first
    print(f"{name:<20} {secs:6.2f}s  {secs / base:5.1f}x  sum={q.first().s:,.2f}")
Output
built-in F.round       1.73s    1.0x  sum=131,294,078.88
Python UDF, pickled    7.51s    4.3x  sum=131,294,078.88
Python UDF, Arrow      3.85s    2.2x  sum=131,294,078.88
pandas UDF             1.81s    1.0x  sum=131,294,078.88
Arrow UDF              1.63s    0.9x  sum=131,294,078.88

All five agree to the cent. On this shared host (load averages up to 9) read ratios: over five runs at one and four million rows, the pickled UDF cost 3.6 to 6.5 times the built-in, the Arrow-optimized one 1.9 to 2.8 times, and the pandas and Arrow UDFs 0.9 to 1.3 times. Arrow transport halves the classic UDF's overhead, but the function still runs four million times; vectorizing removes the calls and, for an expression this cheap, ties the built-in.