Where Warehouses Strained

Why the Data Warehouse Model Strained at Scale

The first analytical platforms copied operational data into one warehouse, cleaned on the way in (schema-on-write). Michael Armbrust, Ali Ghodsi, Reynold Xin and Matei Zaharia summed up what went wrong in their CIDR 2021 paper "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics". On-premises appliances "coupled compute and storage", so a company paid for peak load and peak data at once, which "became very costly as datasets grew"; and more and more data was unstructured (logs, images, text) that a warehouse "could not store and query at all".

The fix was a second tier: land everything in a cheap lake, then load a curated subset into a warehouse such as Redshift 24 or Snowflake. The paper names what that cost. Reliability suffers because two copies of every table are kept consistent by ETL that can fail. Staleness follows, since the warehouse lags the lake, often by days. Advanced analytics is poorly served, because machine-learning libraries want files, not SQL over a proprietary store. The total cost includes storage paid twice, and the curated copy is locked in.

Warehouses are not obsolete: for curated tables used only through SQL, like Analytical SQL and Data Warehouses's marts, they remain the simplest choice. The strain appears when the same data must also feed Spark 129 jobs, model training and streaming, and the copy between lake and warehouse becomes the most fragile job in the platform. The lakehouse deletes that copy.