Apache Parquet 129 is the standard file format of analytical data: it stores each column separately, compressed, with its type and statistics written in the file itself. The conversion step reads the CSV with PyArrow 129 , states the types that matter instead of letting the reader guess, and writes Parquet compressed with Zstandard 126 :
# Step 2: read the CSV with explicit types and write it as a compressed Parquet file.
import os
import pyarrow as pa
import pyarrow.csv as pacsv
import pyarrow.parquet as pq
types = {"id": pa.int32(), "price": pa.decimal128(6, 2), "pages": pa.int16(),
"year": pa.int16(), "in_stock": pa.bool_()}
table = pacsv.read_csv("books.csv", convert_options=pacsv.ConvertOptions(column_types=types))
pq.write_table(table, "books.parquet", compression="zstd")
print(", ".join(f"{f.name}:{f.type}" for f in table.schema if f.type != pa.string()))
print(f"rows={table.num_rows} csv={os.path.getsize('books.csv')} bytes "
f"parquet={os.path.getsize('books.parquet')} bytes")id:int32, price:decimal128(6, 2), pages:int16, year:int16, in_stock:bool rows=6 csv=452 bytes parquet=2666 bytes
The schema now travels with the data: price is an exact decimal, not a floating-point number that only approximates 14.99, and True/False became real Booleans. The other columns (title, author, genre) were inferred as strings, which the listing leaves out.
The size is the surprise: the Parquet file is almost six times larger than the CSV. Parquet spends a fixed overhead on its footer (schema, statistics, offsets) and on headers for every column chunk. On six rows that overhead dominates; on six million rows columnar encoding and compression win easily. Apache Parquet In Depth opens a Parquet file byte by byte, and Compression Codecs Compared compares compression codecs.