Analytical (OLAP) queries scan many rows but few columns. A columnar reader fetches only the named columns (projection pushdown), skips chunks whose statistics rule them out (predicate pushdown, Statistics and Pushdown) and processes each column as a tight array that suits CPU caches and SIMD instructions. This script sums revenue from the CSV export and from a Parquet 129 copy, counting the bytes each read pulls through Linux's /proc/self/io.
import time
import pyarrow.compute as pc, pyarrow.csv as pcsv, pyarrow.parquet as pq
def rchar(): # bytes this process has read so far
return int(open("/proc/self/io").readline().split()[1])
readers = {"CSV, all columns": lambda: pcsv.read_csv("csv-out/orders.csv"),
"Parquet, 1 column": lambda: pq.read_table("orders.parquet", columns=["total"])}
for load in readers.values():
load() # warm up: imports, page cache
for name, load in readers.items():
before, start = rchar(), time.perf_counter()
revenue = pc.sum(load()["total"]).as_py()
print(f"{name:17} read {(rchar() - before) / 1e6:4.2f} MB "
f"in {time.perf_counter() - start:.3f} s, revenue {revenue:,.2f}")CSV, all columns read 6.33 MB in 0.032 s, revenue 3,474,495.41 Parquet, 1 column read 0.19 MB in 0.003 s, revenue 3,474,495.41
The Parquet copy, made with duckdb -c "COPY (FROM read_csv('csv-out/orders.csv')) TO 'orders.parquet'", is 1.48 MB against 6.33 MB. Summing one column read 33 times fewer bytes; it also ran 4 to 45 times faster across six runs on the busy shared 4-CPU host, so trust the byte count more than the clock. Where a warehouse bills by bytes scanned, that ratio is the bill. Platforms keep both layouts: row stores for the application, columnar files for analytics (Analytical SQL and Data Warehouses and Lakehouses, Data Quality and Governance).