All the engines below are open source; the commercial column names the company selling a managed or extended version, not a requirement.
| Engine | Core | Scales out | Best at | Licence | Commercial offering |
|---|---|---|---|---|---|
| Apache Spark 4.2 129 | Scala/JVM | Yes | Big batch, SQL, ML, streaming | Apache-2.0 | Databricks 2,717 , EMR, Fabric and others |
| DuckDB 1.5.6 61,228 | C++ | No | SQL on one machine | MIT | MotherDuck |
| Polars 1.44.2 268,908 | Rust | No (cloud product) | DataFrames on one machine | MIT | Polars Cloud |
| Daft 0.7.25 | Rust | Yes (on Ray) | DataFrames, multimodal data | Apache-2.0 | Eventual |
| Dask 2026.8.0 | Python | Yes | Scaling pandas/NumPy code | BSD-3-Clause | Coiled |
| Ray Data 2.58.0 | C++/Python | Yes | ML preprocessing, inference | Apache-2.0 | Anyscale |
| Flink 2.3.0 129 | Java | Yes | Streaming, unified batch | Apache-2.0 | Confluent 17,245 , Ververica, cloud services |
| Trino 483 403,499 | Java | Yes | Federated, interactive SQL | Apache-2.0 | Starburst 155,289 , Amazon Athena 24 |
Two features separate them in practice: whether the engine needs a cluster process at all (DuckDB, Polars and pandas 16,086 do not), and whether the same code can later run distributed (Spark, Dask, Daft, Ray and Flink can). Daft (github.com/Eventual-Inc/Daft (https://github.com/Eventual-Inc/Daft 5,790 ), pip 21,050 install daft) is the newest entry: a Rust DataFrame engine that runs locally like Polars and distributes over Ray, aimed at multimodal data such as images and documents alongside tables.
"""Daft: the BookNest benchmark. Usage: bench_daft.py T1|T2|T3"""
import sys, time
import daft
from daft import col
D = "/home/dev/v7-l3/ch05/data"
LINES = "order_lines_x10" if sys.argv[1] == "T3" else "order_lines" # T3: 10x rows
t0 = time.perf_counter()
if sys.argv[1] in ("T1", "T3"):
lines = daft.read_parquet(f"{D}/{LINES}.parquet/*.parquet")
books = daft.read_parquet(f"{D}/books.parquet/*.parquet").select(
col("id").cast(daft.DataType.int32()).alias("book_id"), col("genre"))
customers = daft.read_parquet(f"{D}/customers.parquet/*.parquet").select(
col("customer_id"), col("country"))
out = (lines.where(col("status") == "delivered").join(books, on="book_id")
.join(customers, on="customer_id")
.with_column("revenue",
col("qty") * col("unit_price").cast(daft.DataType.float64()))
.with_column("month", col("order_ts").year() * 100 + col("order_ts").month())
.groupby("month", "genre", "country")
.agg(col("revenue").sum(), col("revenue").count().alias("lines"))
.sort(["month", "genre", "country"]))
else:
out = (daft.read_json(f"{D}/raw/orders.jsonl")
.where(col("status") == "delivered").explode("items")
.select((col("order_ts").year() * 100 + col("order_ts").month()).alias("month"),
(col("items").get("qty") * col("items").get("unit_price"))
.alias("revenue"))
.groupby("month").agg(col("revenue").sum()).sort("month"))
res = out.to_pandas()
res.to_parquet(f"/home/dev/v7-l3/bench/out/daft_{sys.argv[1]}.parquet")
print(f"daft {sys.argv[1]} rows={len(res)} revenue={res['revenue'].sum():.2f} "
f"query={time.perf_counter() - t0:.2f}s")Daft's API is close to PySpark 129 's (where, groupby, agg, explode), but version 0.7 had no month truncation, so the listing groups by year * 100 + month, and its JSON reader typed order_ts as a timestamp on its own, where the other engines were given a string. Expect such gaps in young engines and check results, as Benchmark Task and Methodology does.