SELECT ... FROM 'files/orders.parquet' scans JSON, Columnar and Binary Formats's file (four 25,000-order row groups sorted by time, Row Groups and Pages) in place: DuckDB 61,228 reads the footer, then only the needed column chunks of row groups whose statistics can match (Statistics and Pushdown). Linux's /proc/self/io counts the bytes read:
import duckdb, pyarrow
def rchar(): # bytes this process has read so far (Linux)
return int(open("/proc/self/io").read().split()[1])
con = duckdb.connect()
con.execute("SET TimeZone = 'UTC'")
con.sql("SELECT 1").arrow().read_all() # warm up: load the Arrow conversion code
f = "'files/orders.parquet'"
for label, sql in [
("every column", f"SELECT * FROM {f}"),
("items only", f"SELECT items FROM {f}"),
("items, June 2026", f"SELECT items FROM {f} WHERE order_ts >= '2026-06-01'")]:
before = rchar()
rows = con.sql(sql).arrow().read_all().num_rows
print(f"{label:<18}{rows:>7,} rows {rchar() - before:>9,} bytes read")every column 100,000 rows 818,105 bytes read items only 100,000 rows 165,299 bytes read items, June 2026 5,498 rows 106,582 bytes read
The whole 798,714-byte file was read once for SELECT *; one column cost a fifth of that, and the June filter skipped the three row groups whose order_ts maximum is earlier, reading the last group's items and order_ts chunks plus the footer. Set TimeZone: DuckDB reads TIMESTAMPTZ literals in the session zone, which defaults to the machine's (here Asia/Kuala_Lumpur, eight hours off).