| BookNest data | Format | Measured | Why |
|---|---|---|---|
| Partner catalog | XML, JSON | n.a. | Industry vocabularies |
| App order API | JSON + Schema | 253 B/order | Every client parses it |
| Service calls | Protobuf, gRPC | 52 B/order | Generated code |
| Kafka 129 events | Avro 129 + registry | 42 + 5 B | Writer's schema by ID |
| Landing zone | JSONL + zstd 126 | 2.1 MB | Appendable, tolerant |
| Finance exports | CSV | 8.8 MB | Opens in Excel |
| Analytical history | Parquet 129 , zstd | 0.80 MB | Every engine reads it |
| Hive 129 warehouse | ORC, zstd | 0.70 MB | Smallest, Hive-native |
| Engine to engine | Arrow 129 IPC | 11.8 MB | Mapped, never decoded |
Keep text where people and partners read the data and convert to binary on entry. Put a schema'd format and a registry on any stream several teams consume. Store analytics in Parquet unless your one engine favors ORC: Parquet's reader coverage outweighs ORC's 12 percent size advantage on this data.