iterparse yields each element once it ends, but the tree still grows behind it unless you prune it. Three variants on the 100,000-book file show why the usual recipe has two steps:
import resource, sys
from lxml import etree
mode = sys.argv[1]
total = 0.0
for _, book in etree.iterparse("big-catalog.xml", tag="book"): # "end" events by default
total += float(book.findtext("supply/price"))
if mode in ("clear", "clear+delete"):
book.clear() # free this book's children
if mode == "clear+delete":
while book.getprevious() is not None: # and drop the empty shells before it
del book.getparent()[0]
peak = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss >> 10
print(f"{mode:13} total {total:.2f}, peak {peak} MiB")~/de-venv/bin/python ../s2-26/make_big.py 100000
for mode in keep clear clear+delete; do ~/de-venv/bin/python iterparse.py $mode; doneOutput
keep total 2245674.08, peak 680 MiB clear total 2245674.08, peak 30 MiB clear+delete total 2245674.08, peak 18 MiB
clear() empties a book but leaves the empty <book> attached to the root, so memory still grows with the record count, slowly; deleting the preceding siblings keeps it flat for any size. Other rules for big files: pass tag= so the C layer filters events, avoid XPath expressions that look backwards (preceding::), and use huge_tree=True only for trusted input, since it lifts libxml2 3,427 's limits on text nodes and depth.