Arrow IPC and Feather Files

Arrow 129 's IPC format serializes record batches exactly as they sit in memory, with a FlatBuffers message (FlatBuffers) describing the schema and buffer locations. It has two forms: the streaming format, a sequence of messages for sockets and pipes, and the file format, which adds a footer for random access. Feather version 2 is the IPC file format under its old name; .arrow and .feather files are the same thing. Buffers may be compressed with LZ4 or zstd 126 , at the cost of a copy on read:

ipc.py: memory-mapping an Arrow file instead of reading itPython
import os
import pyarrow as pa, pyarrow.compute as pc, pyarrow.feather as feather, pyarrow.ipc as ipc
import pyarrow.parquet as pq
table = pq.read_table("../parquet/orders.parquet")
feather.write_feather(table, "orders.arrow", compression="uncompressed")
feather.write_feather(table, "orders-zstd.arrow", compression="zstd")
for name in ("orders.arrow", "orders-zstd.arrow"):
    before = pa.total_allocated_bytes()
    loaded = ipc.open_file(pa.memory_map(name)).read_all()   # map the file, don't read it
    allocated = pa.total_allocated_bytes() - before
    print(f"{name:18} {os.path.getsize(name):>10,} bytes on disk, {allocated:>10,} bytes "
          f"allocated; revenue {pc.sum(loaded['total']).as_py():,}")
Output
orders.arrow       11,820,354 bytes on disk,          0 bytes allocated; revenue 3,474,495.41
orders-zstd.arrow   2,173,794 bytes on disk, 11,816,384 bytes allocated; revenue 3,474,495.41

Opening the uncompressed file allocated zero bytes: the table's buffers point straight into the memory-mapped file, and the operating system pages in only what the sum touches. The zstd file is 5.4 times smaller but must be decompressed into 11.8 MB of fresh memory. Both are far larger than the 0.8 MB Parquet 129 file, because Arrow keeps values decoded and ready to compute on. The rule of thumb: Parquet for storage and exchange between systems, Arrow IPC for handing data between processes or caching it on a local disk.