With a k8s:// master URL, spark-submit asks the Kubernetes 5,150 API server to create a driver pod; the driver then requests executor pods itself, connects to them, and deletes them when the application ends, leaving the completed driver pod (and its logs) until you remove it. Spark 4.2 129 's documentation requires Kubernetes 1.34 or later, cluster DNS, and a service account that may create pods, services and config maps.
Instead of spark-submit, many teams declare jobs as Kubernetes resources and let an operator submit them, retry them and record their status. Two Apache-licensed operators exist: the community Kubeflow Spark Operator (github.com/kubeflow/spark-operator (https://github.com/kubeflow/spark-operator 3,153 ), v2.5.2 on 31 Jul 2026, SparkApplication and ScheduledSparkApplication resources) and the Spark project's own Apache Spark Kubernetes Operator (github.com/apache/spark-kubernetes-operator (https://github.com/apache/spark-kubernetes-operator 323 ), 1.0.0 on 26 Jul 2026, SparkApplication and SparkCluster, installed with helm install spark spark/spark-kubernetes-operator from https://apache.github.io/spark-kubernetes-operator 126 ). A Kubeflow resource looks like this (from the project's examples; not run here):
apiVersion: sparkoperator.k8s.io/v1beta2
kind: SparkApplication
metadata:
name: spark-pi-python
spec:
type: Python
mode: cluster
image: docker.io/apache/spark:4.0.4
mainApplicationFile: local:///opt/spark/examples/src/main/python/pi.py
sparkVersion: 4.0.4
driver: {cores: 1, memory: 512m, serviceAccount: spark-operator-spark}
executor: {instances: 1, cores: 1, memory: 512m}kubectl 5,150 apply -f creates the job and kubectl get sparkapplications reports its state, so Spark jobs fit into GitOps tooling and schedulers such as Volcano or Apache YuniKorn for gang scheduling and queues.