Infoset and Tree Model

Building the Infoset and Tree Model

The XML Information Set (W3C, Second Edition 2004) defines what a parse means: document, element, attribute, character, comment and processing-instruction information items, with properties such as [children], [parent] and [normalized value]. DOM, XDM and lxml 3,063 's elements are its concrete forms. Getting there involves rewriting the text, which this document shows by mixing Windows line ends, an entity, a declared default, an NMTOKENS attribute and two spellings of one namespace:

infoset.py: what the parser changes between the bytes and the treePython
from lxml import etree
raw = (b'<?xml version="1.0"?>\r\n<!DOCTYPE book [<!ENTITY nest "BookNest">\r\n'
       b'<!ATTLIST book format CDATA "paperback" tags NMTOKENS #IMPLIED>]>\r\n'
       b'<book id="  b1\r\n b2 " tags="  sea\r\n  fiction ">\r\n'
       b'  <title>&nest; &amp; Co</title><!-- note -->\r\n</book>')
root = etree.fromstring(raw, etree.XMLParser(attribute_defaults=True))
print("attributes:", dict(root.attrib))      # line breaks -> spaces; NMTOKENS also trimmed
print("children:", [n.tag if isinstance(n.tag, str) else "comment" for n in root])
print("title text:", repr(root.findtext("title")))  # entity expanded, &amp; decoded
print("text before title:", repr(root.text))         # CR LF became LF
a = etree.fromstring(b'<x:book xmlns:x="urn:booknest"/>')
b = etree.fromstring(b'<book xmlns="urn:booknest"/>')
print("names:", a.tag, b.tag, a.tag == b.tag)       # the name is URI + local name
Output
attributes: {'id': '  b1  b2 ', 'tags': 'sea fiction', 'format': 'paperback'}
children: ['title', 'comment']
title text: 'BookNest & Co'
text before title: '\n  '
names: {urn:booknest}book {urn:booknest}book True

Line ends became \n before anything else; inside attribute values every line break and tab became a space, and because the DTD declares tags as NMTOKENS, its spaces were also collapsed and trimmed (id was not declared, so it kept them). The entity and &amp; are gone from the text, and the declared default format appeared. Namespace prefixes are resolved too: both documents in the last test name the same element, {urn:booknest}book in lxml's notation, whatever the prefix. None of this survives a round trip byte for byte, which is why XML signatures (XML Security) canonicalize first.