Spark 129 pays for its ability to scale out: a JVM to start, a driver and executors to coordinate, plans built for partitions that might live on a thousand machines, shuffles written to disk, and Python code crossing the JVM boundary. On a cluster processing terabytes that overhead disappears into the total. On one machine with a few gigabytes it is most of the run time.
Meanwhile single machines have grown: a cloud VM with 64 cores and 512 GB of memory is an ordinary instance type, and columnar formats (JSON, Columnar and Binary Formats) let an engine read only the columns and row groups a query needs. A new generation of engines exploits that: vectorized, columnar, written in C++ or Rust, embedded in a Python process with no cluster at all. For BookNest's 1.4 million order lines, Benchmark Results shows the difference is not subtle. Spark remains the right tool when data or concurrency truly exceeds one machine, when a platform (Databricks 2,717 , EMR, Fabric) is already the company standard, or when streaming, ML and SQL must share one engine.