MapReduce's Costs

MapReduce's Rigid Stages and Disk I/O Cost

Every MapReduce job has the same shape: read from HDFS, map, sort and spill map output to local disk, shuffle it to the reducers, reduce, and write the result back to HDFS with three replicas. A pipeline of several steps becomes a chain of jobs, and each boundary costs a replicated write, a full read back and fresh task JVMs.

MapReduce writes to disk at every job boundary; Spark pipelines steps in memory between shuffles
MapReduce writes to disk at every job boundary; Spark 129 pipelines steps in memory between shuffles

That hurts iterative algorithms such as k-means, which reread the same data on every pass, and interactive queries, which wait seconds for JVMs to start. The design was sound in 2004, when RAM was scarce and machines failed often: a lost task simply reran from the previous step's files.