Libraries that speak Arrow 129 can hand each other columns by pointer, through the Arrow C data interface (exposed in Python as the __arrow_c_stream__ PyCapsule protocol) or directly through pyarrow 129 . This script converts three order columns into Polars 268,908 and pandas 16,086 and checks whether the order_id values still live at the same memory address, then lets DuckDB 61,228 query the Arrow table in place by its variable name:
import time
import duckdb, pandas as pd, polars as pl, pyarrow as pa, pyarrow.parquet as pq
orders = pq.read_table("../parquet/orders.parquet", columns=["order_id", "channel", "total"])
address = orders["order_id"].chunk(0).buffers()[1].address # where the ids live in memory
def check(label, convert, ids):
start = time.perf_counter()
result = convert()
secs = time.perf_counter() - start
shared = ids(result).chunk(0).buffers()[1].address == address
print(f"{label:24} {secs * 1000:7.1f} ms, same id buffer: {shared}")
check("polars.from_arrow", lambda: pl.from_arrow(orders, rechunk=False),
lambda df: df.to_arrow()["order_id"])
check("to_pandas(ArrowDtype)", lambda: orders.to_pandas(types_mapper=pd.ArrowDtype),
lambda df: pa.chunked_array(pa.array(df["order_id"].array)))
check("to_pandas() to NumPy", lambda: orders.to_pandas(),
lambda df: pa.chunked_array([df["order_id"].to_numpy()]))
revenue = duckdb.sql("SELECT channel, sum(total) FROM orders GROUP BY channel ORDER BY 1")
print("DuckDB on the Arrow table:", ", ".join(f"{c} {r:,}" for c, r in revenue.fetchall()))polars.from_arrow 17.2 ms, same id buffer: True to_pandas(ArrowDtype) 0.9 ms, same id buffer: True to_pandas() to NumPy 207.7 ms, same id buffer: False DuckDB on the Arrow table: android 1,213,501.62, ios 1,557,143.92, web 703,849.87
Polars (1.44.2) and pandas with ArrowDtype keep pointing at Arrow's id buffer. Polars still spent some time because it rewrites string columns into its own string-view layout; pandas with Arrow types did almost nothing. The classic to_pandas() copied: it merged the four row-group chunks into NumPy blocks and turned every DECIMAL into a Python Decimal object, roughly 200 times slower than the Arrow-backed conversion on this shared 4-CPU host (three runs, 158 to 208 ms against about 1 ms). DuckDB found orders as a Python variable, scanned its Arrow buffers without importing them, and produced the same channel revenues as jq 133,477 in Construction and Reduce. In pandas 2.x, pd.read_parquet(path, dtype_backend="pyarrow") gives the Arrow-backed frame directly.