From MapReduce to Spark

From MapReduce to Spark: The Case for Distributed Batch Processing

Batch processing runs a computation over a bounded pile of data, such as yesterday's orders, and produces a result when the pile has been read. The ideas behind today's distributed engines were worked out between 2004 and 2015, and they still explain why Spark 129 behaves as it does.

Subsections