BookNest's data now arrives as XML catalog records, JSON and Parquet 129 order files, Kafka 129 events and dbt 37,942 marts, and an orchestrator moves it every night. This chapter gives it one durable home: a lakehouse, where open Parquet files in object storage behave like warehouse tables that Spark 129 , Trino 403,499 and DuckDB 61,228 can all read and write. You build it on the 4-CPU host with MinIO 30,943 , Apache Iceberg 129 , Delta Lake 237,929 , Trino and DuckDB, then make the data trustworthy and governable, and finish by running BookNest's whole platform end to end.
What you will learn
How object storage, open table formats and catalogs fit together, and how Iceberg and Delta Lake work inside
Running MinIO, Iceberg, Trino and DuckDB locally, and streaming Kafka topics into Iceberg tables
Data quality with GX Core, Soda Core 2,433 and dbt tests, plus anomaly detection, data contracts and lineage
Governance, PII masking, security, lakehouse costs and the managed platforms compared
Sections
- Warehouse to Lakehouse
- Object Storage and MinIO
- Open Table Formats
- Apache Iceberg Tables In Depth
- Iceberg Catalogs
- Iceberg in Production
- Delta Lake and delta-rs
- Hudi and Format Comparison
- Query Engines on the Lakehouse
- Streaming into the Lakehouse
- Data Quality
- Anomalies, Contracts, Lineage
- Governance
- Security and Lakehouse Costs
- Managed Lakehouse Platforms
- The Capstone
- Test Yourself!