The tokenizer is a state machine: in content it looks for < or &; after < it decides between a start tag, an end tag (</), a comment (<!--), CDATA (<![CDATA[), a processing instruction (<?) or a declaration; inside a tag it reads a name, then attribute names, = and quoted values. Each start tag pushes its name on a stack and each end tag must match the top. A conforming parser must stop at the first fatal error (XML 1.0, section 1.2); there is no recovery as in HTML. This loop feeds seven broken snippets to xmllint 3,427 :
check() {
printf '%s' "$2" > t.xml
msg=$(xmllint --noout t.xml 2>&1 | head -n 1 | sed 's/^t.xml:1: //')
printf '%-17s %s\n' "$1" "${msg:-OK}"
}
check "mismatched tags" '<book><title>Salt</book>'
check "two roots" '<book/><book/>'
check "duplicate attr" '<book id="b1" id="b2"/>'
check "unquoted attr" '<book id=b1/>'
check "bare ampersand" '<title>Salt & Saffron</title>'
check "undeclared prefix" '<bn:book/>'
check "bad name" '<1book/>'
check "well-formed" '<book id="b1"><title>Salt & Saffron</title></book>'mismatched tags parser error : Opening and ending tag mismatch: title line 1 and book two roots parser error : Extra content at the end of the document duplicate attr parser error : Attribute id redefined unquoted attr parser error : AttValue: " or ' expected bare ampersand parser error : xmlParseEntityRef: no name undeclared prefix namespace error : Namespace prefix bn on book is not defined bad name parser error : StartTag: invalid element name well-formed OK
Every message names the rule broken: nesting, a single root, unique attributes, quoted values, & starting a reference, declared prefixes (Namespaces in XML, a layer above XML 1.0) and name syntax. Parsers report the first error only, so a broken feed needs a fix-and-rerun loop, or xmllint --recover to salvage what it can for a human to inspect, never for loading.