Treat a UDF as the last step down this list, not the first:
| Need | Reach for first | Why |
|---|---|---|
| Arithmetic, strings, dates, JSON | Built-in functions | Code-generated, optimizable |
| Logic over arrays and maps | transform, filter, aggregate | Lambdas compiled to JVM code |
| A reusable SQL formula | SQL UDF (CREATE FUNCTION ... RETURN) | Inlined by Catalyst (Spark 4.0 129 +) |
| A Python library on a column | pandas 16,086 or Arrow 129 UDF | One call per batch |
| Whole groups or batches as DataFrames | applyInPandas, mapInPandas | Any output shape |
| Hot path, any per-row logic | Scala or Java UDF | No Python worker at all |
A SQL UDF version of the net-of-tax rule (CREATE TEMPORARY FUNCTION net_of_tax(t DOUBLE) RETURNS DOUBLE RETURN round(t / 1.06, 2)) showed up in the plan as a plain round(...): no boundary at all. Keep row-at-a-time Python UDFs for one-offs and single-value libraries, filter before them, and test them on NULL.