Analytical SQL and Data Warehouses ran BookNest's analytics in one database on one machine. Once data outgrows one disk, one memory bank or one nightly window, you need an engine that spreads work over many cores and machines and survives losing any of them. Apache Spark 129 has been that engine for more than a decade.
You follow a BookNest query from Python through Spark's internals, run Spark 4.2 on the WSL2 6 workstation, process a million-row sample order history, read the Spark UI, tune the job, and finally benchmark Spark against Polars 268,908 , DuckDB 61,228 and Dask. Timings come from a shared 4-CPU host, so read ratios rather than seconds.
What you will learn
Why MapReduce gave way to Spark, what Spark 4.x adds, and how Spark runs a job internally.
How to install Spark, use its shells and package a configured spark-submit job.
How to process data with PySpark 129 DataFrames, Spark SQL, joins, windows and UDFs.
How partitioning, caching, the Spark UI and explain() help you fix skew, spills and memory errors.
How to test, deploy and secure Spark jobs, and when to pick another engine.
Sections
- From MapReduce to Spark
- Apache Spark 4.x Today
- Spark Under the Hood
- Running Spark
- DataFrames and Laziness
- Spark SQL and Joins
- UDFs and Arrow
- Partitioning and Caching
- Spark UI and Plans
- Tuning and Memory
- Structured Streaming
- Testing Spark Code
- Spark on Kubernetes
- Managed Spark Platforms
- Securing a Spark Cluster
- Spark Alternatives
- Test Yourself!