MapReduce and Hadoop

Google's MapReduce Paper and Hadoop's Rise

In 2004 Jeffrey Dean and Sanjay Ghemawat published "MapReduce: Simplified Data Processing on Large Clusters." You write map (records to key-value pairs) and reduce (combine the values of one key); the framework splits the input, schedules thousands of machines, moves intermediate data and reruns failures. Google already ran more than a thousand such jobs a day.

Doug Cutting and Mike Cafarella reimplemented the design in Java, and in January 2006 it became Apache Hadoop 129 , named after Cutting's son's toy elephant, with Yahoo! as its biggest user. Hadoop 2.0 (2012) added YARN, letting other engines such as Spark 129 share the cluster. HDFS and YARN live on (Hadoop 3.5.0 shipped in April 2026, and Spark 4.2.0 is built against it); MapReduce as a programming model faded.