The cluster is kind's single node, created with its own kubeconfig so it cannot disturb other contexts, with the lane's working directory mounted into the node so pods can read BookNest's Parquet 129 files. Pulling the official apache/spark:4.2.0-python3 image failed 25 times on this workstation's network (tls: bad record MAC mid-download), so the image was built locally from the installed Spark 129 distribution, following the layout of Spark's own kubernetes/dockerfiles, and loaded with kind load docker-image.
spark.master k8s://https://127.0.0.1:33643
spark.submit.deployMode cluster
spark.kubernetes.container.image l3-spark:4.2.0-python3
spark.kubernetes.authenticate.driver.serviceAccountName spark
spark.kubernetes.driver.request.cores 500m
spark.kubernetes.executor.request.cores 500m
spark.kubernetes.driver.volumes.hostPath.booknest.mount.path /booknest
spark.kubernetes.driver.volumes.hostPath.booknest.mount.readOnly true
spark.kubernetes.driver.volumes.hostPath.booknest.options.path /booknest
spark.kubernetes.executor.volumes.hostPath.booknest.mount.path /booknest
spark.kubernetes.executor.volumes.hostPath.booknest.mount.readOnly true
spark.kubernetes.executor.volumes.hostPath.booknest.options.path /booknestexport KUBECONFIG=/home/dev/v7-l3/k8s/kubeconfig # lane 3's kind cluster, l3-spark
spark-submit --properties-file k8s/booknest-k8s.conf --name booknest-genre \
--conf spark.executor.instances=2 local:///booknest/jobs/k8s_genre.py 2> logs/k8s.log
kubectl get pods
kubectl logs -l spark-role=driver --tail=-1 | grep -A2 "^Technology"NAME READY STATUS RESTARTS AGE booknest-genre-63f1d1a0f91db9ea-driver 0/1 Completed 0 66s Technology 8,878,336.00 Cooking 6,791,664.00 Fiction 6,550,749.92
The job file (jobs/k8s_genre.py, the genre query of Spark Under the Hood with a final collect()) ran unchanged: only the master URL and properties differ from local mode, and the revenue matches Dependencies and --py-files. spark-submit waited and reported the driver pod's phases; a pod listing taken every three seconds showed the two executor pods appear about 20 seconds after submission and vanish when the job ended. Requests of half a core per pod let the three pods fit on the single node. In cluster mode the results live in the driver pod's log, which is why the listing reads it with kubectl 5,150 logs.