Driver and SparkContext

The Driver Program and the SparkContext

The driver is the process that runs your main program. It holds the SparkSession and, inside it, the SparkContext: the connection to the cluster manager, the registry of executors, broadcast variables and accumulators, and the web UI. The driver plans the work, sends tasks out and collects results; it does no heavy lifting unless you pull data into it with collect() or toPandas().

Creating the session and asking the driver about itselfPython
from pyspark.sql import SparkSession
spark = (SparkSession.builder.appName("booknest-engine").master("local[4]")
         .config("spark.ui.port", "33040").getOrCreate())
sc = spark.sparkContext                   # the driver's handle on the cluster
print(sc.master, "|", sc.appName, "|", sc.applicationId)
print("default parallelism:", sc.defaultParallelism, "| UI:", sc.uiWebUrl)
print("driver memory:", spark.conf.get("spark.driver.memory", "1g (default)"))
Output
local[4] | booknest-engine | local-1790853567339
default parallelism: 4 | UI: http://10.255.255.254:33040
driver memory: 1g (default)

If the driver dies, the application dies with it. Keep it lean: the most common driver failure is an out-of-memory error from collecting a large result (Tuning and Memory).