Engine Comparison

Engine Feature and Licence Comparison

All the engines below are open source; the commercial column names the company selling a managed or extended version, not a requirement.

Batch engines compared (versions as of 2 Oct 2026)
Engine Core Scales out Best at Licence Commercial offering
Apache Spark 4.2 129 Scala/JVM Yes Big batch, SQL, ML, streaming Apache-2.0 Databricks 2,717 , EMR, Fabric and others
DuckDB 1.5.6 61,228 C++ No SQL on one machine MIT MotherDuck
Polars 1.44.2 268,908 Rust No (cloud product) DataFrames on one machine MIT Polars Cloud
Daft 0.7.25 Rust Yes (on Ray) DataFrames, multimodal data Apache-2.0 Eventual
Dask 2026.8.0 Python Yes Scaling pandas/NumPy code BSD-3-Clause Coiled
Ray Data 2.58.0 C++/Python Yes ML preprocessing, inference Apache-2.0 Anyscale
Flink 2.3.0 129 Java Yes Streaming, unified batch Apache-2.0 Confluent 17,245 , Ververica, cloud services
Trino 483 403,499 Java Yes Federated, interactive SQL Apache-2.0 Starburst 155,289 , Amazon Athena 24

Two features separate them in practice: whether the engine needs a cluster process at all (DuckDB, Polars and pandas 16,086 do not), and whether the same code can later run distributed (Spark, Dask, Daft, Ray and Flink can). Daft (github.com/Eventual-Inc/Daft (https://github.com/Eventual-Inc/Daft 5,790 ), pip 21,050 install daft) is the newest entry: a Rust DataFrame engine that runs locally like Polars and distributes over Ray, aimed at multimodal data such as images and documents alongside tables.

bench/bench_daft.py: the benchmark in DaftPython
"""Daft: the BookNest benchmark. Usage: bench_daft.py T1|T2|T3"""
import sys, time
import daft
from daft import col
D = "/home/dev/v7-l3/ch05/data"
LINES = "order_lines_x10" if sys.argv[1] == "T3" else "order_lines"   # T3: 10x rows
t0 = time.perf_counter()
if sys.argv[1] in ("T1", "T3"):
    lines = daft.read_parquet(f"{D}/{LINES}.parquet/*.parquet")
    books = daft.read_parquet(f"{D}/books.parquet/*.parquet").select(
        col("id").cast(daft.DataType.int32()).alias("book_id"), col("genre"))
    customers = daft.read_parquet(f"{D}/customers.parquet/*.parquet").select(
        col("customer_id"), col("country"))
    out = (lines.where(col("status") == "delivered").join(books, on="book_id")
           .join(customers, on="customer_id")
           .with_column("revenue",
                        col("qty") * col("unit_price").cast(daft.DataType.float64()))
           .with_column("month", col("order_ts").year() * 100 + col("order_ts").month())
           .groupby("month", "genre", "country")
           .agg(col("revenue").sum(), col("revenue").count().alias("lines"))
           .sort(["month", "genre", "country"]))
else:
    out = (daft.read_json(f"{D}/raw/orders.jsonl")
           .where(col("status") == "delivered").explode("items")
           .select((col("order_ts").year() * 100 + col("order_ts").month()).alias("month"),
                   (col("items").get("qty") * col("items").get("unit_price"))
                   .alias("revenue"))
           .groupby("month").agg(col("revenue").sum()).sort("month"))
res = out.to_pandas()
res.to_parquet(f"/home/dev/v7-l3/bench/out/daft_{sys.argv[1]}.parquet")
print(f"daft {sys.argv[1]} rows={len(res)} revenue={res['revenue'].sum():.2f} "
      f"query={time.perf_counter() - t0:.2f}s")

Daft's API is close to PySpark 129 's (where, groupby, agg, explode), but version 0.7 had no month truncation, so the listing groups by year * 100 + month, and its JSON reader typed order_ts as a timestamp on its own, where the other engines were given a string. Expect such gaps in young engines and check results, as Benchmark Task and Methodology does.