Configuration comes from three places, and the highest wins: values set in code on the session builder, then spark-submit flags (--conf key=value, --driver-memory), then the spark-defaults.conf file in SPARK_CONF_DIR (default $SPARK_HOME/conf).
spark.master local[4]
spark.driver.memory 2g
spark.sql.shuffle.partitions 8
spark.ui.port 33040
spark.eventLog.enabled true
spark.eventLog.dir file:///home/dev/v7-l3/ch05/spark-events
spark.history.fs.logDirectory file:///home/dev/v7-l3/ch05/spark-events
spark.history.ui.port 33180
spark.jars.ivy /home/dev/v7-l3/ch05/ivy
spark.local.dir /home/dev/v7-l3/ch05/tmpexport SPARK_CONF_DIR=$PWD/conf && mkdir -p spark-events
spark-submit jobs/orders_daily.py out/orders_daily # values from the file
spark-submit --conf spark.sql.shuffle.partitions=16 \
jobs/orders_daily.py out/orders_daily # the flag winsOutput
orders_daily: 2,896 rows, 544 days, shuffle partitions 8 orders_daily: 2,896 rows, 544 days, shuffle partitions 16
The file set 8 shuffle partitions; the flag raised it to 16 for one run. The event-log directory must exist: with it missing, pyspark failed at start-up with java.io.FileNotFoundException: File file:/home/dev/v7-l3/ch05/spark-events does not exist. spark-submit --verbose prints every property and its source.