The lake began on Hadoop 129 's HDFS and moved, from about 2015, to cloud object stores such as Amazon S3 24 , Azure 6 Data Lake Storage and Google Cloud 1 Storage. Apache Hive 129 put SQL on top, and its idea of a table was simply a directory: every file under orders/year=2026/month=6/ belonged to the table, directory names encoded partition values (the layout of Globs and Hive Partitioning), and a metastore database remembered schemas and partition lists.
That definition is the technical root of the data swamp. A directory has no transaction boundary, so:
Readers see half-written data: a job writing 40 files is visible after file 1, and a crash leaves debris.
Concurrent writers collide: two jobs overwriting one partition interleave their files.
Listing is the query planner: an S3 listing returns at most 1,000 keys per call, so 100,000 files cost 100 sequential calls before the first byte is read.
Schema lives in two places: the metastore and the files disagree, and a renamed column reads as null.
Updates mean rewrites: deleting one customer's rows rewrites whole partitions by hand.
Teams patched these with conventions (write then rename, _SUCCESS markers). Netflix hit the limits at scale, and in 2017 Ryan Blue and Dan Weeks started Apache Iceberg 129 to replace the directory-as-table model. Uber described Hudi 129 ("Hadoop Upsert Delete and Incremental"), built for the same pressure, in March 2017, and Databricks 2,717 open-sourced Delta Lake 237,929 on 24 April 2019.