Outgrowing One Machine

Why a Single Machine Runs Out of Room

One server hits four walls as data grows. Throughput: reading 10 TB at 500 MB/s takes 5.6 hours, while 100 machines reading a slice each finish in under four minutes. Memory: a sort or join that outgrows RAM spills to disk and slows tenfold. The time window: the nightly job must still finish before the morning reports. Failure: a twelve-hour job that dies at hour eleven starts again from zero.

Scaling out removes the walls but adds coordination and new failure modes. One current server handles jobs that needed a cluster in 2010, which is why DuckDB 61,228 and Polars 268,908 (Spark Alternatives) win many jobs still run on Spark 129 . Scale out when the data, the deadline or the cost of failure leaves no other choice.