Querying Parquet Directly

Querying a Parquet File Without Loading It

SELECT ... FROM 'files/orders.parquet' scans JSON, Columnar and Binary Formats's file (four 25,000-order row groups sorted by time, Row Groups and Pages) in place: DuckDB 61,228 reads the footer, then only the needed column chunks of row groups whose statistics can match (Statistics and Pushdown). Linux's /proc/self/io counts the bytes read:

parquet_scan.py: bytes read for three queries on the same Parquet filePython
import duckdb, pyarrow
def rchar():                                   # bytes this process has read so far (Linux)
    return int(open("/proc/self/io").read().split()[1])
con = duckdb.connect()
con.execute("SET TimeZone = 'UTC'")
con.sql("SELECT 1").arrow().read_all()         # warm up: load the Arrow conversion code
f = "'files/orders.parquet'"
for label, sql in [
    ("every column", f"SELECT * FROM {f}"),
    ("items only", f"SELECT items FROM {f}"),
    ("items, June 2026", f"SELECT items FROM {f} WHERE order_ts >= '2026-06-01'")]:
    before = rchar()
    rows = con.sql(sql).arrow().read_all().num_rows
    print(f"{label:<18}{rows:>7,} rows {rchar() - before:>9,} bytes read")
Output
every column      100,000 rows   818,105 bytes read
items only        100,000 rows   165,299 bytes read
items, June 2026    5,498 rows   106,582 bytes read

The whole 798,714-byte file was read once for SELECT *; one column cost a fifth of that, and the June filter skipped the three row groups whose order_ts maximum is earlier, reading the last group's items and order_ts chunks plus the footer. Set TimeZone: DuckDB reads TIMESTAMPTZ literals in the session zone, which defaults to the machine's (here Asia/Kuala_Lumpur, eight hours off).