The measurements of Python and XML with lxml-XML Parsing Under the Hood on the same 58 MB file, from a shared 4-CPU machine:
| Approach | Peak memory | Notes |
|---|---|---|
| lxml 3,063 tree, ElementTree tree | 681, 580 MiB | XPath and XSLT need a tree |
| JDK DOM | about 880 MB | Includes the JVM |
| xmllint 3,427 tree validation | 699 MB | DTD |
| lxml iterparse, clear and delete | 18 MiB | Constant per record |
| JDK SAX, StAX | 70-95 MB | Mostly the JVM |
| xmllint --stream validation | 21 MB | DTD |
| libxml2 3,427 xmlTextReader (C) | 3.6 MB | Constant |
Trees cost ten to fifteen times the file size, streams a constant. Parsing is CPU-bound (5-10 s per 58 MB here, on a busy machine), so files can arrive compressed at little cost:
~/de-venv/bin/python ../s2-26/make_big.py 100000 && gzip -kf big-catalog.xml
ls -l --block-size=K big-catalog.xml* | awk '{print $5, $9}'
~/de-venv/bin/python gz.pyOutput
58746 big-catalog.xml 657 big-catalog.xml.gz big-catalog.xml 100000 books in 3.65 s big-catalog.xml.gz 100000 books in 3.58 s
The repeated sample data compresses 90-fold, far more than real catalogs. In a pipeline: stream anything that can grow, build a tree per record when you need XPath, run one parser process per file to use all cores, and convert to a columnar format (JSON, Columnar and Binary Formats) once, so later jobs never parse XML again.