On the 100,000-book sample file of lxml Performance (sold-out books relabeled so that StAX reads to the end, and precompiled so that compilation is not timed):
~/de-venv/bin/python ../s2-26/make_big.py 100000
sed -i 's/out-of-stock/in-stock/' big-catalog.xml # so StAX reads to the end
javac -d classes Dom.java Sax.java Stax.java
for api in Dom Sax Stax; do
/usr/bin/time -f " %e s, peak RSS %M KiB" java -Xmx2g -cp classes $api big-catalog.xml
doneOutput
DOM: total 2245674.08, b3 now in-stock 7.93 s, peak RSS 889060 KiB SAX: total 2245674.08 2.89 s, peak RSS 70908 KiB StAX: total 2245674.08 3.46 s, peak RSS 95320 KiB
| Need | DOM | SAX | StAX |
|---|---|---|---|
| Memory | Whole tree (15x the file) | Constant | Constant |
| Random access, edits | Yes | No | No |
| Stop early | After the full parse | Throw an exception | break |
Streaming memory here is mostly the JVM itself; timings come from a shared 4-CPU machine, so compare ratios. A common hybrid streams records with StAX or iterparse and builds a small tree per record.