Batch Processing

Batch Processing and Its Trade-offs

A batch job runs on a schedule or on demand, reads a complete, bounded input, such as one day's orders or one partition of a table, and writes a complete output. Nightly loads into a warehouse, a monthly royalty statement for BookNest's authors and a Spark 129 job that rebuilds a year of sales history are all batch work.

Batch is the default for good reasons. It is simple to reason about: the input is fixed while the job runs, so the same input always gives the same output. It is efficient: reading data in large sequential chunks keeps CPUs and disks busy, and the cluster can be shut down between runs. And it is easy to correct: when you fix a bug, you rerun the affected days, a backfill, and the wrong numbers are replaced.

The cost is latency. A nightly job means dashboards are up to a day old, and a job that fails at 03:00 can leave them two days old. Batch also has a quieter problem at the edges: an order placed at 23:59 that reaches the warehouse at 00:01 may fall into the wrong day unless the job uses event timestamps and rereads recent partitions. Batch Processing with Apache Spark teaches batch processing at scale with Apache Spark, and Orchestration and Pipelines schedules and backfills batch jobs.