Validating While Parsing

Validation is not only a yes or no: a DTD or schema augments the infoset with defaulted and fixed attributes, attribute types (IDs, NMTOKENS) and, for XSD, typed values (the post-schema-validation infoset, or PSVI). Strip the three attributes that booknest.dtd (Document Type Definitions) can supply and let a validating parse restore them:

defaults.sh: a validating parse fills in DTD defaultsShell
# Strip the attributes the DTD can supply, then let a validating parse put them back
sed -e 's/ type="sku"//' -e 's/ seq="1"//' -e 's/ currency="USD"//' \
    -e '1a <!DOCTYPE catalog SYSTEM "booknest.dtd">' booknest-catalog.xml > bare.xml
xmllint --valid --dtdattr bare.xml | grep -m 3 -E '<identifier|<contributor|<price'
grep -m 1 '<contributor' bare.xml
Output
    <identifier type="sku">BN-0001</identifier>
    <contributor role="author" seq="1">
      <price currency="USD">14.99</price>
    <contributor role="author">

A parser that does not read the DTD gives the bare attributes, so the same file means different things to different consumers: give important values explicitly. Validation also runs in a stream, checking each node as it passes (libxml2). On the 100,000-book file with a DOCTYPE and a header added, xmllint 3,427 validates against the DTD in its default tree mode (the loop passes the harmless --nonet there) and with --stream:

stream-valid.sh: DTD validation of a 58 MB file, tree versus streamShell
~/de-venv/bin/python ../s2-26/make_big.py 100000
hdr='<header><sender>S</sender><sent>2026-10-01T08:00:00Z</sent>'
hdr+='<records>100000</records></header>'
sed -i "1s|<catalog>|<!DOCTYPE catalog SYSTEM 'booknest.dtd'>\n<catalog>$hdr|" big-catalog.xml
for flag in --nonet --stream; do
  /usr/bin/time -f "$flag: %e s, peak RSS %M KiB" xmllint --noout --valid $flag big-catalog.xml
done
Output
--nonet: 8.26 s, peak RSS 699012 KiB
--stream: 4.33 s, peak RSS 20980 KiB