The Lakehouse Pattern

The Lakehouse Pattern: Table Semantics on Object Storage

A lakehouse keeps the lake's storage and adds the warehouse's table guarantees with two thin metadata layers.

The five layers of a lakehouse, with the pieces this chapter runs
The five layers of a lakehouse, with the pieces this chapter runs

The bottom two layers are the lake you already know. The table format is the new idea: a table is whatever its latest metadata file says, a list of data files plus the schema and partitioning that describe them. A write creates new data files and a new metadata file, and nothing becomes visible until the catalog swaps its pointer from the old metadata file to the new one in one atomic step. That compare-and-swap gives you ACID commits (readers see a table before or after a write, never halfway), time travel (old metadata files still describe old versions), schema and partition evolution without rewriting files, row-level updates and deletes, and engine independence: Spark 129 writes, Trino 403,499 and DuckDB 61,228 read, and no engine owns the storage.

The price is operational: metadata and small files need maintenance (Iceberg in Production), and the catalog becomes critical infrastructure (Iceberg Catalogs).