etree.parse() returns an ElementTree; elements behave like lists of children with a dictionary of attributes (get(), set()), text and tail strings, and find()/findtext() for simple paths. xpath() runs full XPath 1.0 with variables passed as keyword arguments (no queries built from strings). Every element also knows its parent and source line:
from lxml import etree
tree = etree.parse("booknest-catalog.xml")
root = tree.getroot()
# XPath 1.0 with a variable: results are elements, strings or numbers
for book in root.xpath("book[supply/price >= $min]", min=20):
print(book.get("id"), book.findtext("title"), book.find(".//price").text)
print("total:", root.xpath("sum(//price)"), "| b6 starts on line", root[6].sourceline)
# Elements know their parent, so edits are one call each
b3 = root.xpath("//book[@id = $id]", id="b3")[0]
b3.find("supply/availability").text = "in-stock"
etree.SubElement(b3, "format", kind="hardcover")
print(etree.tostring(b3[-1]), b3[-1].getparent().get("id"), len(b3))
try:
etree.fromstring("<book><title>Unclosed</book>")
except etree.XMLSyntaxError as e:
print("XMLSyntaxError:", e.msg)Output
b2 Patterns of the Deep Web 39.50 b3 Salt and Saffron 24.00 b6 Gardens in Glass 21.30 total: 134.74 | b6 starts on line 99 b'<format kind="hardcover"/>' b3 10 XMLSyntaxError: Opening and ending tag mismatch: title line 1 and book, line 1, column 29
root[6] is b6 because root[0] is the header. xpath() returns a float for sum() and a list for node-sets; namespaced documents need a namespaces={"bn": "..."} mapping (Namespaces and xml: Attributes).