Every step of the run is code you have already met, unchanged or pointed at the platform namespace:
| Chapter | Piece reused | Its job in the platform | Step |
|---|---|---|---|
| 2 XML | booknest-catalog.xml, catalog.xsd, XQuery | Validate and flatten the catalog | catalog |
| 3 Formats | generate_orders.py, MANIFEST.sha256 | Produce the canonical sample data | generate |
| 4 Analytical SQL | dbt 37,942 models and tests, the mart.sales grain | Model and test the mart | build_mart |
| 5 Spark 129 | A local PySpark 129 job, Apache Iceberg Tables In Depth's tables | Land books and orders in Iceberg 129 | load_lake |
| 6 Kafka 129 | Topic booknest.order-events, Kafka Connect 129 | Carry the event stream | stream_events |
| 7 Orchestration | An Airflow 3 129 DAG, dags test | Run the steps, stop on failure | the DAG |
| 8 Lakehouse | MinIO 30,943 , Iceberg, Trino 403,499 , contract, Presidio | Store, gate, query, publish | all steps |
Two pieces changed on the way in. Analytical SQL and Data Warehouses's dbt models targeted PostgreSQL 1,289 and DuckDB 61,228 ; here they run through dbt-trino 1.10.5, so Trino writes them as Iceberg tables (dbt Tests Revisited found that dbt-duckdb could not materialize a table in the Iceberg catalog). Orchestration and Pipelines's pipeline loaded days into PostgreSQL; the capstone keeps its contract, idempotent steps and a check before publishing, and swaps the targets for lake tables. The topic name and key, the table shapes, the contract and the PII classifier did not change: a platform is mostly agreements between pieces, and these held.