Benchmark Setup

Benchmark Setup: BookNest's Order History on a 4-CPU Host

Everything ran on the chapter's workstation: an Intel Core i5-6500 (4 cores, 3.2 GHz) with 32 GB of RAM, WSL2 6 Ubuntu 26.04 225 on a local SSD, and Python 3.14.4. The engines were installed in a separate virtual environment with pip 21,050 install polars duckdb pyarrow 129 pandas 16,086 "dask[dataframe]" daft, which resolved to DuckDB 1.5.6 61,228 , Polars 1.44.2 268,908 , Daft 0.7.25, Dask 2026.8.0, pandas 3.0.6 and PyArrow 25.0.1; Spark 4.2.0 129 ran in local mode on Java 21 with a 4 GB driver. The data is the chapter's sample order history: 1,380,694 order lines and 50,000 customers in Spark-written Parquet 129 , and the 238 MB orders.jsonl they came from. A tenfold copy of the order lines tests how each engine scales:

bench/make_x10.sh: 13.8 million order lines for task T3Shell
#!/bin/bash
# T3 input: the order lines repeated 10 times (13.8 M rows), written by DuckDB.
cd /home/dev/v7-l3/ch05/data && rm -rf order_lines_x10.parquet
TZ=UTC /home/dev/v7-l3/bench-venv/bin/python - <<'PY'
import duckdb
con = duckdb.connect(); con.execute("SET TimeZone='UTC'; SET threads=4")
con.execute("""COPY (SELECT l.* FROM 'order_lines.parquet/*.parquet' l, range(10))
               TO 'order_lines_x10.parquet' (FORMAT parquet, PER_THREAD_OUTPUT true)""")
print(con.execute("""SELECT count(*), typeof(any_value(order_ts))
                      FROM 'order_lines_x10.parquet/*.parquet'""").fetchone())
PY
du -sh order_lines_x10.parquet
Output
(13806940, 'TIMESTAMP')
48M   order_lines_x10.parquet

The machine was shared with two other workloads, whose load average varied between about 1.7 and 5 during the runs, so every result below is a median of three runs and is read as a ratio.