A Decision Table

A Decision Table for BookNest's Data

BookNest's data flows; sizes measured in JSON Lines-Compression Codecs Compared (100,000 orders, or order 1)
BookNest data Format Measured Why
Partner catalog XML, JSON n.a. Industry vocabularies
App order API JSON + Schema 253 B/order Every client parses it
Service calls Protobuf, gRPC 52 B/order Generated code
Kafka 129 events Avro 129 + registry 42 + 5 B Writer's schema by ID
Landing zone JSONL + zstd 126 2.1 MB Appendable, tolerant
Finance exports CSV 8.8 MB Opens in Excel
Analytical history Parquet 129 , zstd 0.80 MB Every engine reads it
Hive 129 warehouse ORC, zstd 0.70 MB Smallest, Hive-native
Engine to engine Arrow 129 IPC 11.8 MB Mapped, never decoded

Keep text where people and partners read the data and convert to binary on entry. Put a schema'd format and a registry on any stream several teams consume. Store analytics in Parquet unless your one engine favors ORC: Parquet's reader coverage outweighs ORC's 12 percent size advantage on this data.