Parquet 129 stores only flat leaf columns, so a nested field such as an order's items list is shredded: each leaf (items.list.element.book_id and its siblings) becomes its own column, and two small integers per value let a reader rebuild the structure, as the Dremel paper describes. The definition level says how many optional or repeated ancestors of a value are actually present, so it encodes nulls and empty lists; the repetition level says at which repeated ancestor a new value starts, where 0 means a new record. The schema fixes each column's maximum levels:
python levels.py | grep -E "items|coupon|^order_id"order_id INT64 max_rep=0 max_def=0 items.list.element.book_id INT32 max_rep=1 max_def=1 items.list.element.qty INT32 max_rep=1 max_def=1 items.list.element.unit_price INT32 max_rep=1 max_def=1 coupon BYTE_ARRAY max_rep=0 max_def=1
levels.py prints each column's maximum levels from pyarrow 129 's ParquetFile(...).schema; a required flat column needs none. These are the levels for orders 1, 2 and 16 (the first with a coupon):
| Order | book_id values | Repetition levels | Definition levels | coupon (definition level) |
|---|---|---|---|---|
| 1 | 3, 5 | 0, 1 | 1, 1 | null (0) |
| 2 | 3, 6, 5 | 0, 1, 1 | 1, 1, 1 | null (0) |
| 16 | 1 | 0 | 1 | READMORE15 (1) |
Repetition level 1 means "another element of the same list" and 0 starts the next order; definition level 0 on coupon marks a null, so no value is stored. With the list and its elements required, only the repeated list counts toward book_id's definition level, and an empty list would be one entry at level 0. With Arrow 129 's default (everything nullable) the maximum would be 4. The levels are RLE-encoded and nearly free for regular data, and a query on book_id reads only that leaf column.