By the late 2010s many companies ran both a lake and a warehouse loaded from it, so every number lived twice, with a pipeline between them that could fall behind. The lakehouse removes the second copy. Databricks 2,717 ' founders named the pattern in a CIDR 2021 paper, characterizing it by open direct-access formats such as Parquet 129 , first-class support for machine learning, and state-of-the-art performance.
The trick is a table format: a layer of metadata over plain Parquet files that records which files make up each version of a table. Writers commit a new version atomically by adding a log entry, so readers never see a half-written state. That one idea brings ACID transactions, updates, deletes and time travel to the lake. Delta Lake 237,929 , Apache Iceberg 129 and Apache Hudi 129 are the three open table formats (Open Table Formats to Hudi and Format Comparison). You can see the mechanism with the deltalake Python package, which writes Delta tables without Spark 129 :
# A lakehouse table is Parquet files plus a transaction log (delta-rs, no Spark needed).
import json, os, re, shutil
import pyarrow as pa
from deltalake import DeltaTable, write_deltalake
shutil.rmtree("books_delta", ignore_errors=True)
books = json.load(open("books.json"))["books"]
rows = pa.Table.from_pylist([{k: b[k] for k in ("id", "title", "price")} for b in books])
write_deltalake("books_delta", rows) # version 0: six books
DeltaTable("books_delta").update(updates={"price": "19.99"}, predicate="id = 1") # version 1
for root, _, files in sorted(os.walk("books_delta")):
for name in sorted(files):
print(os.path.join(root, re.sub(r"[0-9a-f]{8}-[0-9a-f-]{27}", "<uuid>", name)))
log = open("books_delta/_delta_log/00000000000000000001.json").read().splitlines()
print("commit 1 actions:", [next(iter(json.loads(line))) for line in log])
for v in (0, 1):
t = DeltaTable("books_delta", version=v).to_pyarrow_table().sort_by("id")
print(f"version {v}: book 1 costs", t.column("price")[0].as_py())books_delta/part-00000-<uuid>-c000.snappy.parquet books_delta/part-00000-<uuid>-c000.snappy.parquet books_delta/_delta_log/00000000000000000000.json books_delta/_delta_log/00000000000000000001.json commit 1 actions: ['commitInfo', 'add', 'remove'] version 0: book 1 costs 14.99 version 1: book 1 costs 19.99
Parquet files are immutable, so changing one price wrote a whole new data file; commit 1 adds it and removes the old one from the current version without deleting it from disk. That is why version 0 can still be read, and why lakehouse tables need periodic cleanup of files no version uses any more (Iceberg in Production).
| Property | Warehouse | Lake | Lakehouse |
|---|---|---|---|
| Data types | Structured tables | Anything | Anything, tables for structured data |
| Storage format | Vendor's own | Open files | Open files plus open table format |
| Transactions | Yes | No | Yes |
| Engines | The vendor's | Any | Any that support the format |
| Main risk | Lock-in, cost | Swamp | Operational complexity |
The lakehouse is now the default design for new platforms, and Snowflake, BigQuery 1 and Redshift 24 can all query Iceberg tables. Lakehouses, Data Quality and Governance builds BookNest's lakehouse with Iceberg and queries it from Trino 403,499 and DuckDB 61,228 .