At a Larger Scale

What Would Change at a Larger Scale

The capstone rebuilds everything in minutes because BookNest's history is 100,000 orders. The platform's shape survives growth; most deployments do not:

BookNest's platform on one host and at scale
Layer On this host At a hundred times the data
Orchestration airflow dags test, one process Scheduled Airflow 129 on Kubernetes 5,150 , or managed (7.16)
Events One broker, console replay Three or more brokers, RF 3, services producing (6.2)
Batch Spark 129 local[2] Spark on Kubernetes or a managed service (5.13, 5.14)
Storage, catalog MinIO 30,943 on one disk, REST fixture Cloud object storage, an authenticated catalog (Iceberg Catalogs)
Query One Trino 403,499 JVM, 3 GB A Trino cluster with autoscaled workers

Four changes go deeper. Incremental processing replaces rebuilds: Spark MERGEs each new day into its tables (Spark as a Lakehouse Engine) and dbt 37,942 's incremental models (Transforming Data with dbt) touch only recent partitions. Maintenance becomes a job of its own, because a streaming sink commits small files all day (Compaction and Expiry and Orphan Files and Manifests). Write-audit-publish moves to branches, so audits run on isolated data (Branches and Tags). And gates turn statistical: you check each new snapshot exactly and the whole table by sampling and the robust z-scores of Anomaly Detection. Each source also gets an owning team, and its contract becomes the interface between teams.