Test Yourself!

These ten questions cover the behaviors of XML tools that most often surprise working engineers: predicates that apply per parent, untyped values that become doubles or strings, comparisons that fail, rounding rules, token matching, pending updates, source-order numbering, path dialects and attribute normalization. Every snippet ran on the book's workstation with Saxon-HE 12.10 726,956 , BaseX 12.4 693,830 , lxml 6.1.3 3,063 and Python 3.14. Write down what each one prints and why, and only then check Appendix I, which gives the real output, the reason and the section it comes from.

What to look for

Each question pairs two expressions that look interchangeable, so the interesting part is why the two halves differ. Ask of every value whether it is a node, an untyped atomic value, a string, a decimal or a double, since most of the surprises in this chapter came from a value silently changing type (XPath 3.1 Data Types and XPath 3.1 Path Expressions). Ask which processor runs the code: XPath 1.0 in lxml and browsers, XPath 3.1 in Saxon and BaseX, ElementPath in find(). And for the BaseX questions, remember that full-text matching works on tokens and that updates wait on the pending update list until the query ends. Questions 1-5 run in one Saxon call, so a warning on standard error appears before their results.

Checking your answers

check.sh: run the four question filesShell
# Run the four question files; compare with Appendix I
xquery -s:booknest-catalog.xml questions-1-5.xq
basex questions-6-7.xq; echo
xslt3 -s:booknest-catalog.xml -xsl:question-8.xsl
~/de-venv/bin/python questions-9-10.py

Questions

questions-1-5.xq: paths, numbers, comparisons and rounding in Saxon
(: Run with Saxon-HE: xquery -s:booknest-catalog.xml questions-1-5.xq :)
declare namespace err = "http://www.w3.org/2005/xqt-errors";
(: 1. How many items does each path select? (Section 2.6.2) :)
"1: " || count(//contributor/name[1]) || " vs " || count((//contributor/name)[1]),
(: 2. Two sums of the same untyped values. What do they print? (Sections 2.5.3, 2.7.2) :)
let $p := (<p>0.1</p>, <p>0.2</p>)
return "2: " || sum($p) || " vs " || sum($p ! xs:decimal(.)),
(: 3. A value comparison and a general comparison on untyped pages. (Section 2.18.1) :)
"3: " || (try { count(//book[pages gt 400]) } catch * { $err:code }) || " vs "
      || count(//book[pages > 400]),
(: 4. Two ways to show 18.75 plus 10% with two decimals. (Section 2.19.1) :)
"4: " || format-number(20.625, "0.00") || " vs " || round(20.625, 2),
(: 5. Which number is greater? (Sections 2.5.1 and 2.6.3) :)
"5: " || (<n>9</n> > <n>10</n>) || " vs " || (9 > 10)
questions-6-7.xq: full text and the update facility in BaseX
(: Run with BaseX: basex questions-6-7.xq :)
(: 6. Character matching versus token matching. (Section 2.17.1) :)
"6: " || contains("Salt and Saffron", "saffron") || " vs "
      || ("Salt and Saffron" contains text "saffron"),
(: 7. Which price does the <old> element record? (Section 2.17.3) :)
copy $b := <book><price>24.00</price></book>
modify (replace value of node $b/price with "26.00",
        insert node <old>{ string($b/price) }</old> into $b)
return "7: " || serialize($b)
question-8.xsl: numbering in XSLTXML
<!-- 8. Run with Saxon: xslt3 -s:booknest-catalog.xml -xsl:question-8.xsl
     What does xsl:number print for the in-stock books? (Section 2.18.2) -->
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
  <xsl:output method="text"/>
  <xsl:template match="/catalog">
    <xsl:text>8:</xsl:text>
    <xsl:for-each select="book[supply/availability = 'in-stock']">
      <xsl:text> </xsl:text><xsl:number/>
    </xsl:for-each>
    <xsl:text>&#10;</xsl:text>
  </xsl:template>
</xsl:stylesheet>
questions-9-10.py: lxml paths and attribute values in PythonPython
# Run with the venv of Section 2.26: python questions-9-10.py
import xml.etree.ElementTree as ET
from lxml import etree
root = etree.parse("booknest-catalog.xml").getroot()
# 9. The same path through lxml's find() and through xpath(). (Section 2.26.4)
try:
    found = len(root.findall("book[supply/price > 20]"))
except SyntaxError as e:
    found = f"SyntaxError: {e}"
print("9:", found, "vs", len(root.xpath("book[supply/price > 20]")))
# 10. What value does the parser report for this attribute? (Section 2.29.2)
print("10:", repr(ET.fromstring(b'<book id="b1\r\n\tb2"/>').get("id")))