JSON, Columnar and Binary Formats's 100,000 orders, rewritten as XML with lxml 3,063 's streaming xmlfile writer:
import gzip, json
from lxml import etree
src = "/mnt/d/Books/Data Engineering/demos/ch03/data/orders.jsonl" # Section 3.1
keys = ("order_id", "customer_id", "order_ts", "channel", "status")
with etree.xmlfile("orders.xml", encoding="utf-8") as xf, xf.element("orders"):
for line in open(src): # stream: one <order> at a time
o = json.loads(line)
with xf.element("order", {k: str(o[k]) for k in keys}):
for it in o["items"]:
xf.write(etree.Element("item", {k: str(v) for k, v in it.items()}))
total = etree.Element("total", currency=o["currency"])
total.text = str(o["total"])
xf.write(total)
for name in ("orders.xml", src):
data = open(name, "rb").read()
gz = len(gzip.compress(data))
print(f"{name[-12:]:12} {len(data) / 1e6:4.1f} MB, gzip {gz / 1e6:3.1f} MB")Output
orders.xml 21.4 MB, gzip 1.5 MB orders.jsonl 23.6 MB, gzip 1.5 MB
Each order became <order order_id="1" ... status="delivered"> with item and total children. With values in attributes the XML is smaller than the JSON, and gzip makes them equal (parsing both takes a few seconds). The real differences lie elsewhere:
| Aspect | XML | JSON |
|---|---|---|
| Types | Text, typed by a schema | Strings, numbers, Booleans, null |
| Schemas and business rules | XSD, RELAX NG, Schematron | JSON Schema (JSON Schema 2020-12) |
| Browser and API ecosystem | Shrinking | Dominant |
Keep XML where a standard defines it (ONIX, XBRL, invoices, SAML) and convert at the edge (JSON, Columnar and Binary Formats).