Data Lakes and Swamps

The Data Lake Era and the Data Swamp Problem

The lake began on Hadoop 129 's HDFS and moved, from about 2015, to cloud object stores such as Amazon S3 24 , Azure 6 Data Lake Storage and Google Cloud 1 Storage. Apache Hive 129 put SQL on top, and its idea of a table was simply a directory: every file under orders/year=2026/month=6/ belonged to the table, directory names encoded partition values (the layout of Globs and Hive Partitioning), and a metastore database remembered schemas and partition lists.

That definition is the technical root of the data swamp. A directory has no transaction boundary, so:

Teams patched these with conventions (write then rename, _SUCCESS markers). Netflix hit the limits at scale, and in 2017 Ryan Blue and Dan Weeks started Apache Iceberg 129 to replace the directory-as-table model. Uber described Hudi 129 ("Hadoop Upsert Delete and Incremental"), built for the same pressure, in March 2017, and Databricks 2,717 open-sourced Delta Lake 237,929 on 24 April 2019.