Apache Spark

Batch Processing with Apache Spark

Analytical SQL and Data Warehouses ran BookNest's analytics in one database on one machine. Once data outgrows one disk, one memory bank or one nightly window, you need an engine that spreads work over many cores and machines and survives losing any of them. Apache Spark 129 has been that engine for more than a decade.

You follow a BookNest query from Python through Spark's internals, run Spark 4.2 on the WSL2 6 workstation, process a million-row sample order history, read the Spark UI, tune the job, and finally benchmark Spark against Polars 268,908 , DuckDB 61,228 and Dask. Timings come from a shared 4-CPU host, so read ratios rather than seconds.

What you will learn

Sections