Storage

Storage: Where Data Rests Between Stages

Every stage after generation reads from storage and writes back to it, so storage choices ripple through the whole platform. Three families cover most needs today. Object storage (Amazon S3 24 , Azure 6 Data Lake Storage, Google Cloud 1 Storage, or an S3-compatible store on your own servers) holds files of any size cheaply and durably, and is the foundation of data lakes. Databases hold rows you query and update with low latency. Analytical warehouses (Analytical SQL and Data Warehouses) store data in columns for fast scans and aggregates. Open table formats such as Apache Iceberg 129 and Delta Lake 237,929 (Lakehouses, Data Quality and Governance) now let plain files in object storage behave like warehouse tables, with transactions and time travel.

Storage families and their trade-offs
Storage Good at Weak at Example
Object storage Cheap, durable bulk files Updating single records S3, MinIO 30,943
Transactional database Fast reads and writes of rows Scanning billions of rows PostgreSQL 1,289
Analytical warehouse Aggregates over huge tables High-rate single-row updates BigQuery 1 , DuckDB 61,228
Event log Ordered, replayable streams Ad hoc queries Kafka 129
Lakehouse table Warehouse features on open files Operational simplicity Iceberg, Delta

Two further decisions matter as much as the product. The file format decides how fast data can be read back: a columnar format such as Parquet 129 lets a query read only the columns it needs, which is why JSON, Columnar and Binary Formats spends several sections on formats. The temperature of data, how often it is read, decides what you pay: recent orders are hot and queried all day, while five-year-old click logs are cold and belong in a cheaper storage class with a retention rule that eventually deletes them. Storage is usually the cheapest part of a platform; scanning it badly is the expensive part.