Memory and Throughput

Memory and Throughput Trade-offs in Pipelines

The measurements of Python and XML with lxml-XML Parsing Under the Hood on the same 58 MB file, from a shared 4-CPU machine:

Peak memory for one 58 MB feed
Approach Peak memory Notes
lxml 3,063 tree, ElementTree tree 681, 580 MiB XPath and XSLT need a tree
JDK DOM about 880 MB Includes the JVM
xmllint 3,427 tree validation 699 MB DTD
lxml iterparse, clear and delete 18 MiB Constant per record
JDK SAX, StAX 70-95 MB Mostly the JVM
xmllint --stream validation 21 MB DTD
libxml2 3,427 xmlTextReader (C) 3.6 MB Constant

Trees cost ten to fifteen times the file size, streams a constant. Parsing is CPU-bound (5-10 s per 58 MB here, on a busy machine), so files can arrive compressed at little cost:

Reading the same file plain and gzip-compressed
~/de-venv/bin/python ../s2-26/make_big.py 100000 && gzip -kf big-catalog.xml
ls -l --block-size=K big-catalog.xml* | awk '{print $5, $9}'
~/de-venv/bin/python gz.py
Output
58746 big-catalog.xml
657 big-catalog.xml.gz
big-catalog.xml     100000 books in 3.65 s
big-catalog.xml.gz  100000 books in 3.58 s

The repeated sample data compresses 90-fold, far more than real catalogs. In a pipeline: stream anything that can grow, build a tree per record when you need XPath, run one parser process per file to use all cores, and convert to a columnar format (JSON, Columnar and Binary Formats) once, so later jobs never parse XML again.