Every MapReduce job has the same shape: read from HDFS, map, sort and spill map output to local disk, shuffle it to the reducers, reduce, and write the result back to HDFS with three replicas. A pipeline of several steps becomes a chain of jobs, and each boundary costs a replicated write, a full read back and fresh task JVMs.

That hurts iterative algorithms such as k-means, which reread the same data on every pass, and interactive queries, which wait seconds for JVMs to start. The design was sound in 2004, when RAM was scarce and machines failed often: a lost task simply reran from the previous step's files.