Parquet Metadata in Practice

Reading and Writing Parquet Metadata for Real

Everything above came from the footer, and every serious engine exposes it. In Python, pyarrow.parquet.read_metadata(path) returns the FileMetaData (created_by, row groups, schema) and .row_group(i).column(j) a chunk's encodings, codec, sizes and statistics, as stats.py used. DuckDB 61,228 turns the same footer into tables:

metadata.sh: the footer as an SQL table in DuckDBShell
duckdb -c "
SELECT row_group_id AS rg, path_in_schema AS col, encodings, stats_min_value AS min,
       stats_max_value AS max, total_compressed_size AS bytes
FROM parquet_metadata('orders.parquet')
WHERE path_in_schema IN ('status', 'total') AND row_group_id IN (0, 3);"
Output
┌───────┬─────────┬────────────────────────────┬───────────┬──────────┬───────┐
│  rg   │   col   │         encodings          │    min    │   max    │ bytes │
│ int64 │ varchar │          varchar           │  varchar  │ varchar  │ int64 │
├───────┼─────────┼────────────────────────────┼───────────┼──────────┼───────┤
│     0 │ status  │ PLAIN, RLE, RLE_DICTIONARY │ cancelled │ returned │  3022 │
│     0 │ total   │ PLAIN, RLE, RLE_DICTIONARY │ 12.74     │ 199.60   │ 24299 │
│     3 │ status  │ PLAIN, RLE, RLE_DICTIONARY │ cancelled │ shipped  │  3313 │
│     3 │ total   │ PLAIN, RLE, RLE_DICTIONARY │ 12.74     │ 205.49   │ 23840 │
└───────┴─────────┴────────────────────────────┴───────────┴──────────┴───────┘

PLAIN in the encodings list is the dictionary page. String statistics compare bytes, so status runs from "cancelled" to "returned", except in the last group, the only one still holding shipped orders. DuckDB shows total as decimals because the footer records the DECIMAL(9,2) logical type over INT32.

On the write side, parquet_kv_metadata('orders.parquet') lists two key-value entries: booknest.source (20 bytes), which write_parquet.py attached with replace_schema_metadata, and ARROW:schema (1,208 bytes), the serialized Arrow 129 schema pyarrow 129 adds so types such as time zones round-trip exactly. created_by (parquet-cpp-arrow version 25.0.1) lets readers work around known bugs of old writers. DuckDB also has parquet_schema() and parquet_file_metadata(). Checking metadata after any new writer or setting is the cheapest tuning habit there is.