The Full Pipeline

The Full Pipeline from Catalog to Lakehouse

BookNest's data starts in three places: the catalog service's XML feed, the order service's JSON Lines export, and the events of each order's life. The figure shows the route each takes into the lakehouse and out again.

BookNest's data platform: sources, ingestion, the Iceberg lakehouse, the dbt mart and the gates
BookNest's data platform: sources, ingestion, the Iceberg 129 lakehouse, the dbt 37,942 mart and the gates

The catalog feed is validated against XML and Its Toolchain's XSD and flattened by an XQuery into six JSON lines. The orders come from JSON, Columnar and Binary Formats's seeded generator, checked against its checksums, and a Spark 129 job writes books, orders and order lines as Iceberg tables in a fresh namespace, platform. The events travel as a live service would send them: onto Apache Kafka and Managed Cloud Kafka's topic booknest.order-events, keyed by order_id, and through Kafka Connect 129 's Iceberg sink (Streaming into the Lakehouse) into platform.order_events. From then on every step reads tables, not files.

dbt models the mart in Trino 403,499 : each order's status as its latest event, Analytical SQL and Data Warehouses's one-row-per-line sales fact, and Orchestration and Pipelines's daily gross per genre. Gates sit at every boundary, and the last step publishes by moving an Iceberg tag, so readers see the previous good mart or the new one, never a half-built one. Every step is idempotent and reports one line of evidence a person can check.