Classic PySpark 129 starts a JVM inside your Python process, so client and cluster share one Spark 129 version and one dependency set. Spark Connect 129 splits them: the client sends unresolved logical plans as Protocol Buffers 59,573 over gRPC to a Connect server, which runs them and streams results back as Apache Arrow 129 batches. Start one with sbin/start-connect-server.sh (default port 15002; 33015 here) and install the 1.6 MB, JVM-free pyspark-client package in its own virtual environment.
from pyspark.sql import SparkSession # from pyspark-client: no JVM in this process
spark = SparkSession.builder.remote("sc://localhost:33015").getOrCreate()
print(type(spark).__module__, spark.version)
spark.sql("""FROM parquet.`data/orders.parquet`
|> WHERE status = 'delivered'
|> AGGREGATE count(*) AS orders, sum(total) AS revenue GROUP BY channel
|> ORDER BY revenue DESC""").show()pyspark.sql.connect.session 4.2.0 +-------+------+-----------+ |channel|orders| revenue| +-------+------+-----------+ | ios|401094|13939528.77| |android|311750|10850818.70| | web|177859| 6189935.90| +-------+------+-----------+
Each |> step of Spark 4.0's SQL pipe syntax transforms the previous result. Connect has no SparkContext or RDD API, so df.rdd code must be rewritten first. pyspark-client 4.2.0 pulled in pandas 3.0.6 16,086 , which PySpark warns it "does not yet fully support"; pin pandas<3 beside it.