spark.executor.memory (default 1g, minimum 450m) sets each executor's heap, spark.driver.memory (1g) the driver's, and spark.executor.memoryOverhead defaults to 10% of the heap with a 384 MiB minimum (memoryOverheadFactor is 0.4 for non-JVM jobs on Kubernetes 5,150 ). Heap sizes must be set before the JVM starts, with --driver-memory or spark-defaults.conf: setting spark.driver.memory in the session builder of a running client-mode driver is silently too late. The listing ran under spark-submit --driver-memory 512m, 1g and 4g; in local mode the driver's heap is also the executor's.
"""Print what --driver-memory turns into: spark-submit --driver-memory 1g l5102_memory.py"""
import json, urllib.request
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("memory").getOrCreate()
sc = spark.sparkContext
heap = sc._jvm.java.lang.Runtime.getRuntime().maxMemory() / 2**20
url = f"{sc.uiWebUrl}/api/v1/applications/{sc.applicationId}/executors"
unified = json.load(urllib.request.urlopen(url))[0]["maxMemory"] / 2**20 # the unified pool
print(f"driver memory {sc.getConf().get('spark.driver.memory'):<4} heap {heap:6.0f} MiB "
f"unified {unified:6.1f} MiB (heap - 300) x 0.6 = {(heap - 300) * 0.6:6.1f} MiB")driver memory 512m heap 512 MiB unified 127.2 MiB (heap - 300) x 0.6 = 127.2 MiB driver memory 1g heap 1024 MiB unified 434.4 MiB (heap - 300) x 0.6 = 434.4 MiB driver memory 4g heap 4096 MiB unified 2277.6 MiB (heap - 300) x 0.6 = 2277.6 MiB
The formula holds exactly, and shows why small heaps hurt: 512 MiB leaves 127 MiB for all execution and storage. On clusters, 4 or 5 cores per executor is a common starting point; huge heaps lengthen GC pauses.