The driver is the process that runs your main program. It holds the SparkSession and, inside it, the SparkContext: the connection to the cluster manager, the registry of executors, broadcast variables and accumulators, and the web UI. The driver plans the work, sends tasks out and collects results; it does no heavy lifting unless you pull data into it with collect() or toPandas().
from pyspark.sql import SparkSession
spark = (SparkSession.builder.appName("booknest-engine").master("local[4]")
.config("spark.ui.port", "33040").getOrCreate())
sc = spark.sparkContext # the driver's handle on the cluster
print(sc.master, "|", sc.appName, "|", sc.applicationId)
print("default parallelism:", sc.defaultParallelism, "| UI:", sc.uiWebUrl)
print("driver memory:", spark.conf.get("spark.driver.memory", "1g (default)"))Output
local[4] | booknest-engine | local-1790853567339 default parallelism: 4 | UI: http://10.255.255.254:33040 driver memory: 1g (default)
If the driver dies, the application dies with it. Keep it lean: the most common driver failure is an out-of-memory error from collecting a large result (Tuning and Memory).