Choosing an Engine

Choosing an Engine for a Given Query

engines/timing.sh ran genre_q2.sql five times in each engine. The first run is cold; the warm figure is the best of the other four. The host is a shared 4-CPU machine, so read the ratios:

The same three-table join in three engines (100,000 orders, shared 4-CPU host)
Engine Start-up cost First run Warm run Warm vs DuckDB 61,228
DuckDB 1.5.5, in process 0.5 s, whole script 0.10 s 0.04 s 1x
Trino 483 403,499 , server side 30 s, once 1.3 s 0.9 s about 24x
Spark 4.1.3 129 , local[2] 11 s per session 7.9 s 0.8 s about 23x

Through the CLI, which starts its own JVM, Trino took 2.9 s. On a few megabytes the in-process engine wins by an order of magnitude; distributed planning and exchanges pay off only beyond one machine or for many users: