Choosing a Spark Platform

Choosing a Platform for BookNest's Workload

BookNest's nightly jobs read a few million rows and finish in minutes on one workstation, so the honest first answer is that BookNest does not need managed Spark 129 yet: a single VM, or DuckDB 61,228 or Polars 268,908 (Spark Alternatives), runs the same work for less. The decision changes as the order history grows past one machine. Then follow where the data and the team already are: EMR Serverless for a company on AWS 24 with data in S3, Managed Service for Apache Spark serverless on Google Cloud 1 next to BigQuery 1 , Fabric for a Microsoft 365 and Power BI shop, and Databricks 2,717 when a lakehouse with governance, notebooks and ML on any cloud justifies a platform fee. Keep the code portable: the jobs in this chapter use only open-source Spark APIs and Parquet 129 , so they run on all five, and spark-submit arguments, not code, carry the platform-specific settings.