lxml Performance

Performance and libxml2 Bindings

lxml 3,063 's parsing happens in C; Python objects are created only for the elements you touch, through Cython proxies. The cost that matters in a pipeline is memory: a tree of a 58 MB file takes more than ten times that. iterparse() fixes that by yielding each element as it ends; calling clear() on it frees the subtree. make_big.py repeats the six books 100,000 times (sample data), and bench.py sums every price with a tree and with iterparse:

bench.py: whole tree versus iterparse, in lxml and ElementTreePython
import resource, sys, time
import xml.etree.ElementTree as ET
from lxml import etree
def tree(mod):          # whole document in memory
    return sum(float(p.text) for p in mod.parse("big-catalog.xml").getroot().iter("price"))
def stream(mod):        # iterparse, discarding each book once read
    total = 0.0
    for _, book in mod.iterparse("big-catalog.xml", events=("end",), tag="book") \
            if mod is etree else mod.iterparse("big-catalog.xml"):
        if book.tag == "book":
            total += float(book.find("supply/price").text)
            book.clear()
    return total
mod = {"ET": ET, "lxml": etree}[sys.argv[1]]
t = time.perf_counter()
total = {"tree": tree, "stream": stream}[sys.argv[2]](mod)
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
secs = time.perf_counter() - t
print(f"{sys.argv[1]:4} {sys.argv[2]:6} {secs:5.2f} s {peak:6.0f} MiB  {total:.2f}")
Building a 100,000-book file and running the four combinations
~/de-venv/bin/python make_big.py 100000 && echo "$(du -m big-catalog.xml | cut -f1) MB"
for m in lxml ET; do for k in tree stream; do ~/de-venv/bin/python bench.py $m $k; done; done
Output
58 MB
lxml tree    3.99 s    681 MiB  2245674.08
lxml stream  5.41 s     31 MiB  2245674.08
ET   tree    6.34 s    579 MiB  2245674.08
ET   stream  6.45 s     27 MiB  2245674.08

Streaming cut peak memory about twentyfold in both libraries. The timings come from a shared 4-CPU machine busy with other jobs (they varied by 2x between runs), so compare ratios, not seconds: plain parsing is close, and lxml's lead shows in XPath, validation and XSLT, which run in C (more on streaming in Streaming Large Documents).