Data Lakes and Swamps

The Data Lake and the Data Swamp

James Dixon, then chief technology officer of Pentaho, coined the term in a blog post of 14 October 2010: "If you think of a datamart as a store of bottled water – cleansed and packaged and structured for easy consumption – the data lake is a large body of water in a more natural state." A data lake keeps raw data of every kind, in its original form, on cheap storage, and leaves the structure to whoever reads it (schema-on-read, Schema-on-Write vs on-Read). The first lakes ran on Hadoop 129 's HDFS; today they are object-storage buckets holding CSV, JSON, Parquet 129 , images and logs side by side.

Lakes are cheap (Amazon S3 24 Standard lists $0.023 per gigabyte a month for the first 50 TB in us-east-1), any engine can read the files, and a new source can land today and be modeled next month. Without discipline, though, a lake turns into a data swamp: files nobody can explain, conflicting copies of the same table, no owner to ask, and no way to tell whether a file is complete or half-written by a job that crashed.

The common cure is to divide the lake into zones by quality. Raw data lands untouched in a bronze zone, cleaned and conformed copies live in silver, and business-ready tables in gold, a layout Databricks 2,717 calls the medallion architecture. Add a catalog of what each dataset means and who owns it (Governance) and quality checks between zones (Data Quality), and the swamp stays a lake.