lxml 3,063 's parsing happens in C; Python objects are created only for the elements you touch, through Cython proxies. The cost that matters in a pipeline is memory: a tree of a 58 MB file takes more than ten times that. iterparse() fixes that by yielding each element as it ends; calling clear() on it frees the subtree. make_big.py repeats the six books 100,000 times (sample data), and bench.py sums every price with a tree and with iterparse:
import resource, sys, time
import xml.etree.ElementTree as ET
from lxml import etree
def tree(mod): # whole document in memory
return sum(float(p.text) for p in mod.parse("big-catalog.xml").getroot().iter("price"))
def stream(mod): # iterparse, discarding each book once read
total = 0.0
for _, book in mod.iterparse("big-catalog.xml", events=("end",), tag="book") \
if mod is etree else mod.iterparse("big-catalog.xml"):
if book.tag == "book":
total += float(book.find("supply/price").text)
book.clear()
return total
mod = {"ET": ET, "lxml": etree}[sys.argv[1]]
t = time.perf_counter()
total = {"tree": tree, "stream": stream}[sys.argv[2]](mod)
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
secs = time.perf_counter() - t
print(f"{sys.argv[1]:4} {sys.argv[2]:6} {secs:5.2f} s {peak:6.0f} MiB {total:.2f}")~/de-venv/bin/python make_big.py 100000 && echo "$(du -m big-catalog.xml | cut -f1) MB"
for m in lxml ET; do for k in tree stream; do ~/de-venv/bin/python bench.py $m $k; done; done58 MB lxml tree 3.99 s 681 MiB 2245674.08 lxml stream 5.41 s 31 MiB 2245674.08 ET tree 6.34 s 579 MiB 2245674.08 ET stream 6.45 s 27 MiB 2245674.08
Streaming cut peak memory about twentyfold in both libraries. The timings come from a shared 4-CPU machine busy with other jobs (they varied by 2x between runs), so compare ratios, not seconds: plain parsing is close, and lxml's lead shows in XPath, validation and XSLT, which run in C (more on streaming in Streaming Large Documents).