XML usually enters a data platform at the edge, as files from outside, and leaves it soon after as rows. The stages in between are where this chapter's tools work:

Spark 129 (built-in XML source since 4.0, spark.read.option("rowTag", "book").xml(path)) and pandas 16,086 (read_xml()) can do the last hop, flattening records into rows, but they cannot validate a feed or reshape a nested vocabulary, so the XML-native stages stay. Install them once: xmllint 3,427 and xsltproc from libxml2 3,427 , a JDK (Saxon Extension Functions compiles Java), and Saxon-HE 12.10 726,956 , Saxonica's open-source engine for XPath 3.1, XQuery 3.1 and XSLT 3.0, with two small wrappers the chapter uses throughout:
sudo apt-get install -y libxml2-utils xsltproc openjdk-21-jdk-headless
mkdir -p ~/xml-tools/saxon ~/bin && cd ~/xml-tools/saxon
v=SaxonHE12-10
curl -fsSLO --retry 3 https://github.com/Saxonica/Saxon-HE/releases/download/$v/${v}J.zip
unzip -oq ${v}J.zip
cat > ~/bin/xquery <<'EOF'
#!/bin/sh
java -cp "$HOME/xml-tools/saxon/saxon-he-12.10.jar" net.sf.saxon.Query \
'!method=adaptive' '!omit-xml-declaration=yes' "$@" && echo
EOF
cat > ~/bin/xslt3 <<'EOF'
#!/bin/sh
exec java -cp "$HOME/xml-tools/saxon/saxon-he-12.10.jar" net.sf.saxon.Transform "$@"
EOF
chmod +x ~/bin/xquery ~/bin/xslt3xquery runs XPath and XQuery (a superset of XPath): -qs: takes the expression, -s: names the input, and the adaptive output method prints one item per line with strings quoted, so types stay visible. xslt3 runs stylesheets (XSLT Flow and Construction). Ubuntu 225 adds ~/bin to PATH at your next login. If unzip complains about a missing "central directory", the download arrived truncated: run curl 3,008 again.