VTD-XML's Non-Extractive Model

VTD-XML 7,086 (Virtual Token Descriptor; GPL-2.0, last release 2.13.4 in 2017) keeps the document's bytes and records each token as a 64-bit entry of offset, length, type and depth instead of creating objects. XPath runs over that array, and edits splice new bytes into the original without re-parsing. Run it with java -cp ~/xml-tools/vtd/vtd-xml-2.13.4.jar Vtd.java:

Vtd.java: XPath over token records and an in-place updateJava
import com.ximpleware.*;
public class Vtd {
  public static void main(String[] args) throws Exception {
    VTDGen gen = new VTDGen();
    gen.parseFile("booknest-catalog.xml", false);     // records tokens as offsets, no objects
    VTDNav nav = gen.getNav();
    System.out.println(nav.getTokenCount() + " tokens, " + nav.getXML().length() + " bytes");
    AutoPilot ap = new AutoPilot(nav);
    ap.selectXPath("//book[supply/availability = 'out-of-stock']/supply/price/text()");
    int i = ap.evalXPath();                            // i is a token index, not a node
    System.out.println("token " + i + " at " + nav.getTokenOffset(i) + ": " + nav.toString(i));
    XMLModifier mod = new XMLModifier(nav);
    mod.updateToken(i, "26.00");                       // splice new bytes, no re-parse
    mod.output("repriced.xml");
    gen.parseFile("repriced.xml", false);
    System.out.println("repriced.xml, same token: " + gen.getNav().toString(i));
  }
}
Output
238 tokens, 3858 bytes
token 126 at 2031: 24.00
repriced.xml, same token: 26.00

Re-parsing the output puts 26.00 at the same token index: only bytes changed. The index costs roughly a third to a half of the file size, far less than a DOM, but there is no DOM or SAX API, no external entities, only XPath 1.0, a GPL licence and no release in nine years: an idea worth knowing rather than a dependency to add.


VTD XML

VTD-XML 7,086 (Virtual Token Descriptor for XML) is a group of cross-platform, non-extractive processing methods for XML. In traditional, extractive parsing, a lexical analyzer represents each token as a discrete string object copied out of the source. VTD-XML instead keeps the source document intact in memory and describes each token by its offset and length within that source - an approach its developers describe as "document-centric".

Because tokens are recorded as compact numeric records rather than materialized as separate string/object instances, VTD-XML skips the object-oriented modeling step used by parsers such as DOM, eliminating the associated object-creation and garbage-collection costs. This lets navigation stay fast even on large documents, and supports incremental updates - modifying a document in place without a full re-parse. The tradeoffs are that the in-memory representation increases the effective document size by roughly 30% to 50%, and VTD-XML is not compatible with DOM, SAX, external DTD entities, or certain validation techniques.

Navigation in VTD-XML is performed through a cursor object (commonly named VTDNav in the Java and C/C# implementations), with an AutoPilot helper class used to run XPath-like expressions over the indexed document.

Full nameVirtual Token Descriptor for XML
Processing modelNon-extractive; tokens recorded as offset/length pairs instead of extracted strings
Source documentKept intact in memory alongside the token index
Memory overheadApproximately 30% to 50% larger than the raw document size
CompatibilityNot compatible with DOM, SAX, external DTD entities, or certain validation techniques
Typical usesFast random access and incremental (in-place) updates without a full re-parse
ch09-vtdxml-nav.javaJava
VTDGen vg = new VTDGen();
if (vg.parseFile("books.xml", false)) {
    VTDNav vn = vg.getNav();
    AutoPilot ap = new AutoPilot(vn);
    ap.selectXPath("//book/@id");

    int i;
    while ((i = ap.evalXPath()) != -1) {
        System.out.println("Book id: " + vn.toString(i));
    }
}