The History Server

The History Server for Completed Applications

The live UI disappears with its driver. With spark.eventLog.enabled, the driver also writes every event to a log, and the History Server (Logging and History Server) replays those logs into the same UI and the same REST API, which makes after-the-fact checks scriptable.

Querying a finished job through the History Server's REST APIShell
export SPARK_CONF_DIR=$PWD/conf SPARK_LOG_DIR=$PWD/logs SPARK_PID_DIR=$PWD/pids
$SPARK_HOME/sbin/start-history-server.sh > /dev/null
spark-submit jobs/orders_daily.py out/orders_daily        # writes an event log, then exits
sleep 15                                                  # the server rescans every 10 s
H=http://localhost:33180/api/v1/applications
APP=$(curl -s "$H?status=completed&limit=1" | jq -r '.[0].id')
curl -s "$H/$APP" | jq -c '{name, attempts: [.attempts[] | {duration, completed}]}'
curl -s "$H/$APP/stages" | jq -r 'sort_by(-.executorRunTime) | .[0:3][]
  | "stage \(.stageId): \(.numTasks) tasks, \(.executorRunTime) ms of task time"'
$SPARK_HOME/sbin/stop-history-server.sh > /dev/null
Output
orders_daily: 2,896 rows, 544 days, shuffle partitions 8
{"name":"orders_daily","attempts":[{"duration":42241,"completed":true}]}
stage 3: 4 tasks, 18008 ms of task time
stage 0: 1 tasks, 1530 ms of task time
stage 2: 1 tasks, 1281 ms of task time

The nightly job's scan-join-aggregate stage took nearly all the task time, which is where tuning effort belongs. Event logs grow with every job: in production write them to shared storage (HDFS, S3, ADLS), enable spark.eventLog.rolling.enabled for long applications and spark.history.fs.cleaner.enabled (default off; maxAge 7 days) to expire old ones. A REST check like this one can alert when a stage's task time jumps from one night to the next.