The Lakehouse

By the late 2010s many companies ran both a lake and a warehouse loaded from it, so every number lived twice, with a pipeline between them that could fall behind. The lakehouse removes the second copy. Databricks 2,717 ' founders named the pattern in a CIDR 2021 paper, characterizing it by open direct-access formats such as Parquet 129 , first-class support for machine learning, and state-of-the-art performance.

The trick is a table format: a layer of metadata over plain Parquet files that records which files make up each version of a table. Writers commit a new version atomically by adding a log entry, so readers never see a half-written state. That one idea brings ACID transactions, updates, deletes and time travel to the lake. Delta Lake 237,929 , Apache Iceberg 129 and Apache Hudi 129 are the three open table formats (Open Table Formats to Hudi and Format Comparison). You can see the mechanism with the deltalake Python package, which writes Delta tables without Spark 129 :

A Delta table is Parquet files plus a log of commitsPython
# A lakehouse table is Parquet files plus a transaction log (delta-rs, no Spark needed).
import json, os, re, shutil
import pyarrow as pa
from deltalake import DeltaTable, write_deltalake
shutil.rmtree("books_delta", ignore_errors=True)
books = json.load(open("books.json"))["books"]
rows = pa.Table.from_pylist([{k: b[k] for k in ("id", "title", "price")} for b in books])
write_deltalake("books_delta", rows)                       # version 0: six books
DeltaTable("books_delta").update(updates={"price": "19.99"}, predicate="id = 1")  # version 1
for root, _, files in sorted(os.walk("books_delta")):
    for name in sorted(files):
        print(os.path.join(root, re.sub(r"[0-9a-f]{8}-[0-9a-f-]{27}", "<uuid>", name)))
log = open("books_delta/_delta_log/00000000000000000001.json").read().splitlines()
print("commit 1 actions:", [next(iter(json.loads(line))) for line in log])
for v in (0, 1):
    t = DeltaTable("books_delta", version=v).to_pyarrow_table().sort_by("id")
    print(f"version {v}: book 1 costs", t.column("price")[0].as_py())
Output
books_delta/part-00000-<uuid>-c000.snappy.parquet
books_delta/part-00000-<uuid>-c000.snappy.parquet
books_delta/_delta_log/00000000000000000000.json
books_delta/_delta_log/00000000000000000001.json
commit 1 actions: ['commitInfo', 'add', 'remove']
version 0: book 1 costs 14.99
version 1: book 1 costs 19.99

Parquet files are immutable, so changing one price wrote a whole new data file; commit 1 adds it and removes the old one from the current version without deleting it from disk. That is why version 0 can still be read, and why lakehouse tables need periodic cleanup of files no version uses any more (Iceberg in Production).

Warehouse, lake and lakehouse compared
Property Warehouse Lake Lakehouse
Data types Structured tables Anything Anything, tables for structured data
Storage format Vendor's own Open files Open files plus open table format
Transactions Yes No Yes
Engines The vendor's Any Any that support the format
Main risk Lock-in, cost Swamp Operational complexity

The lakehouse is now the default design for new platforms, and Snowflake, BigQuery 1 and Redshift 24 can all query Iceberg tables. Lakehouses, Data Quality and Governance builds BookNest's lakehouse with Iceberg and queries it from Trino 403,499 and DuckDB 61,228 .