XML and Its Toolchain gave BookNest's catalog an XML vocabulary. The rest of a data platform speaks other formats: orders leave the app as JSON, land as JSON Lines, reach finance as CSV, cross Kafka 129 as Avro 129 or Protocol Buffers 59,573 and settle into Parquet 129 . Each hop is a format decision whose trade-offs you live with for years. Every format here writes the same 100,000 BookNest sample orders (Order Events as JSON Lines), reused in Analytical SQL and Data Warehouses, Batch Processing with Apache Spark, Apache Kafka and Managed Cloud Kafka and Lakehouses, Data Quality and Governance.
What you will learn:
Where JSON loses data silently, and how JSON Lines streams it
JSON Schema, JSONPath, jq 133,477 and yq 16,045 ; CSV that survives real pipelines
Row versus columnar layout; Avro, Protocol Buffers, MessagePack 229,447 and CBOR
Parquet, ORC and Arrow 129 internals, codecs, format choice and safe parsing
Sections
- JSON's Data Model
- JSON Lines
- JSON Schema 2020-12
- Querying JSON with JSONPath
- jq in Depth
- yq for JSON and YAML
- CSV Done Right
- Row vs Columnar Layout
- Apache Avro
- Protocol Buffers
- Thrift and Zero-Copy Formats
- MessagePack and CBOR
- Apache Parquet In Depth
- Apache ORC
- Apache Arrow
- Compression Codecs Compared
- Schema Evolution
- Choosing a Data Format
- Data Format Security
- Test Yourself!