Executor and Driver Memory

Configuring Executor and Driver Memory

spark.executor.memory (default 1g, minimum 450m) sets each executor's heap, spark.driver.memory (1g) the driver's, and spark.executor.memoryOverhead defaults to 10% of the heap with a 384 MiB minimum (memoryOverheadFactor is 0.4 for non-JVM jobs on Kubernetes 5,150 ). Heap sizes must be set before the JVM starts, with --driver-memory or spark-defaults.conf: setting spark.driver.memory in the session builder of a running client-mode driver is silently too late. The listing ran under spark-submit --driver-memory 512m, 1g and 4g; in local mode the driver's heap is also the executor's.

What --driver-memory turns intoPython
"""Print what --driver-memory turns into: spark-submit --driver-memory 1g l5102_memory.py"""
import json, urllib.request
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("memory").getOrCreate()
sc = spark.sparkContext
heap = sc._jvm.java.lang.Runtime.getRuntime().maxMemory() / 2**20
url = f"{sc.uiWebUrl}/api/v1/applications/{sc.applicationId}/executors"
unified = json.load(urllib.request.urlopen(url))[0]["maxMemory"] / 2**20   # the unified pool
print(f"driver memory {sc.getConf().get('spark.driver.memory'):<4}  heap {heap:6.0f} MiB  "
      f"unified {unified:6.1f} MiB  (heap - 300) x 0.6 = {(heap - 300) * 0.6:6.1f} MiB")
Output
driver memory 512m  heap    512 MiB  unified  127.2 MiB  (heap - 300) x 0.6 =  127.2 MiB
driver memory 1g    heap   1024 MiB  unified  434.4 MiB  (heap - 300) x 0.6 =  434.4 MiB
driver memory 4g    heap   4096 MiB  unified 2277.6 MiB  (heap - 300) x 0.6 = 2277.6 MiB

The formula holds exactly, and shows why small heaps hurt: 512 MiB leaves 127 MiB for all execution and storage. On clusters, 4 or 5 cores per executor is a common starting point; huge heaps lengthen GC pauses.