Spark 4.0 129 also brought SQL functions, session variables, collations, an XML source and Python UDTFs. Spark 4.1 (December 2025) added Declarative Pipelines, a real-time Structured Streaming mode, Arrow-native Python UDFs and generally available SQL scripting. Spark 4.2 (July 2026) added GEOMETRY and GEOGRAPHY types, a CHANGES clause for change data capture, metric views, Java 25 support, Arrow-optimized Python UDFs by default (UDFs and Arrow) and a web UI with dark mode and zoomable plans (Spark UI and Plans).
QUALIFY filters on a window function without the subquery Analytical SQL and Data Warehouses needed: QUALIFY rank() OVER (PARTITION BY yr ORDER BY copies DESC) = 1 names The Quiet Harbor the best seller of both 2025 and 2026 in the sample data. SQL scripting runs BEGIN ... END blocks with DECLAREd variables.