Statistics and Pushdown

Statistics and Predicate Pushdown

The footer stores minimum, maximum and null count for every column chunk, and the optional page index (written here with write_page_index=True) stores the same per page. A reader evaluates a query's WHERE clause against them and skips every row group or page that cannot match: predicate pushdown.

stats.py: which row groups can a filter skip?Python
from datetime import datetime, timezone
import pyarrow.dataset as ds, pyarrow.parquet as pq
meta = pq.read_metadata("orders.parquet")
for i in range(meta.num_row_groups):
    ts, cust = meta.row_group(i).column(2).statistics, meta.row_group(i).column(1).statistics
    print(f"row group {i}: order_ts {ts.min:%Y-%m-%d} to {ts.max:%Y-%m-%d}, "
          f"customer_id {cust.min} to {cust.max}")
fragment = next(ds.dataset("orders.parquet").get_fragments())
june = datetime(2026, 6, 1, tzinfo=timezone.utc)
for label, condition in (("order_ts >= 2026-06-01", ds.field("order_ts") >= june),
                         ("customer_id = 42", ds.field("customer_id") == 42)):
    kept = fragment.split_by_row_group(filter=condition)   # groups the stats can't rule out
    print(f"{label:23} reads {len(kept)} of {meta.num_row_groups} row groups")
Output
row group 0: order_ts 2025-01-01 to 2025-05-17, customer_id 1 to 5000
row group 1: order_ts 2025-05-17 to 2025-10-01, customer_id 1 to 5000
row group 2: order_ts 2025-10-01 to 2026-02-14, customer_id 1 to 5000
row group 3: order_ts 2026-02-14 to 2026-06-30, customer_id 1 to 5000
order_ts >= 2026-06-01  reads 1 of 4 row groups
customer_id = 42        reads 4 of 4 row groups

Statistics help only when values cluster. The orders are in time order, so each group covers a disjoint date range and a June query reads one group; customers are spread over all groups. Lakehouse writers therefore sort or cluster data on filtered columns (Lakehouses, Data Quality and Governance), and sorting_columns records the sort in the footer. For equality lookups on unsorted, high-cardinality columns, Parquet 129 also has optional per-chunk bloom filters, which pyarrow 25 129 writes with bloom_filter_options.