Streaming Large Documents

Streaming Very Large XML Documents

iterparse yields each element once it ends, but the tree still grows behind it unless you prune it. Three variants on the 100,000-book file show why the usual recipe has two steps:

iterparse.py: keep everything, clear each book, or clear and deletePython
import resource, sys
from lxml import etree
mode = sys.argv[1]
total = 0.0
for _, book in etree.iterparse("big-catalog.xml", tag="book"):    # "end" events by default
    total += float(book.findtext("supply/price"))
    if mode in ("clear", "clear+delete"):
        book.clear()                                 # free this book's children
    if mode == "clear+delete":
        while book.getprevious() is not None:        # and drop the empty shells before it
            del book.getparent()[0]
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss >> 10
print(f"{mode:13} total {total:.2f}, peak {peak} MiB")
Running the three variants
~/de-venv/bin/python ../s2-26/make_big.py 100000
for mode in keep clear clear+delete; do ~/de-venv/bin/python iterparse.py $mode; done
Output
keep          total 2245674.08, peak 680 MiB
clear         total 2245674.08, peak 30 MiB
clear+delete  total 2245674.08, peak 18 MiB

clear() empties a book but leaves the empty <book> attached to the root, so memory still grows with the record count, slowly; deleting the preceding siblings keeps it flat for any size. Other rules for big files: pass tag= so the C layer filters events, avoid XPath expressions that look backwards (preceding::), and use huge_tree=True only for trusted input, since it lifts libxml2 3,427 's limits on text nodes and depth.