Every earlier chapter feeds the lakehouse. XML and Its Toolchain's catalog XML and JSON, Columnar and Binary Formats's order, customer and event files land unchanged in a bucket called booknest-raw (the bronze zone of Data Lakes and Swamps) and become Iceberg 129 tables (books, orders, order_items) in the namespace booknest, stored in a bucket called warehouse. Analytical SQL and Data Warehouses's dbt 37,942 models and tests run against those tables (Data Quality), Batch Processing with Apache Spark's Spark 129 becomes one lakehouse engine among several, Apache Kafka and Managed Cloud Kafka's order events stream in through Kafka Connect 129 (Streaming into the Lakehouse), and Orchestration and Pipelines's pipeline and data contract supply the schedule, contracts and lineage of Anomalies, Contracts, Lineage.
The orders are JSON, Columnar and Binary Formats's generated sample data (100,000 orders, seed 7), and XML and Its Toolchain and Analytical SQL and Data Warehouses both computed their non-cancelled gross revenue as 3,303,427.30. Read the figure straight from JSON, Columnar and Binary Formats's Parquet 129 file with DuckDB 61,228 first; the Iceberg tables of Apache Iceberg Tables In Depth must reproduce it exactly.
SET TimeZone = 'UTC';
SELECT count(*) AS orders,
sum(len(items)) AS order_lines,
sum(total) AS net_revenue,
sum(list_sum(list_transform(items, lambda i: i.qty * i.unit_price)))
FILTER (WHERE status <> 'cancelled') AS non_cancelled_gross
FROM '/mnt/d/Books/Data Engineering/demos/ch03/out/orders.parquet';┌────────┬─────────────┬───────────────┬─────────────────────┐ │ orders │ order_lines │ net_revenue │ non_cancelled_gross │ │ int64 │ int128 │ decimal(38,2) │ decimal(38,2) │ ├────────┼─────────────┼───────────────┼─────────────────────┤ │ 100000 │ 137944 │ 3474495.41 │ 3303427.30 │ └────────┴─────────────┴───────────────┴─────────────────────┘
Net revenue (after coupons) and gross revenue (quantity times list price) differ on purpose, so a reconciliation must compare like with like. The host runs at UTC+8, so every engine in this chapter is pinned to UTC (SET TimeZone, spark.sql.session.timeZone, Trino 403,499 's SET TIME ZONE) and dates agree across all three.