Ask the questions in order and stop at the first clear answer:
| Question | If yes |
|---|---|
| Does the data fit comfortably on one large machine (up to a few hundred GB per job)? | DuckDB 61,228 for SQL, Polars 268,908 for DataFrames |
| Is the team's code already pandas 16,086 or NumPy that must scale out? | Dask |
| Is the work ML preprocessing or batch inference on GPUs? | Ray Data (or Daft on Ray) |
| Is streaming the main workload, with some batch? | Flink 129 , or Spark 129 Structured Streaming |
| Must one query join many live systems in place? | Trino 403,499 |
| Is the data multi-terabyte, or the platform already Spark-based? | Spark, usually managed (Managed Spark Platforms) |
For BookNest today, the nightly jobs belong in DuckDB or Polars: they finish in under a second, need no cluster, and run inside an Airflow 129 task (Orchestration and Pipelines). The chapter's Spark skills still pay off, because the concepts carry over (lazy plans, pushdown, partitioning, skew, explain), and because the jobs were written against Parquet 129 and SQL, moving one to Spark when the order history outgrows a machine is a rewrite of a few lines, not a migration.