BookNest's orders, catalog feed and app events land as raw files. Spark 129 is the batch engine in the middle: it removes duplicates and bad records, joins orders to the catalog and customers, and writes clean, aggregated tables for the sales mart (Analytical SQL and Data Warehouses) and the lakehouse (Lakehouses, Data Quality and Governance), while Airflow 129 (Orchestration and Pipelines) decides when each job runs.
At a million orders DuckDB 61,228 could do the same work (Spark Alternatives). Spark earns its place because the same code runs on a laptop and on a hundred nodes, Lakehouses, Data Quality and Governance's table formats treat it as their reference engine, and one engine covers SQL, Python and streaming. This chapter builds the nightly orders_daily job: revenue and copies sold per book, genre and day.