VTD-XML 7,086 (Virtual Token Descriptor; GPL-2.0, last release 2.13.4 in 2017) keeps the document's bytes and records each token as a 64-bit entry of offset, length, type and depth instead of creating objects. XPath runs over that array, and edits splice new bytes into the original without re-parsing. Run it with java -cp ~/xml-tools/vtd/vtd-xml-2.13.4.jar Vtd.java:
import com.ximpleware.*;
public class Vtd {
public static void main(String[] args) throws Exception {
VTDGen gen = new VTDGen();
gen.parseFile("booknest-catalog.xml", false); // records tokens as offsets, no objects
VTDNav nav = gen.getNav();
System.out.println(nav.getTokenCount() + " tokens, " + nav.getXML().length() + " bytes");
AutoPilot ap = new AutoPilot(nav);
ap.selectXPath("//book[supply/availability = 'out-of-stock']/supply/price/text()");
int i = ap.evalXPath(); // i is a token index, not a node
System.out.println("token " + i + " at " + nav.getTokenOffset(i) + ": " + nav.toString(i));
XMLModifier mod = new XMLModifier(nav);
mod.updateToken(i, "26.00"); // splice new bytes, no re-parse
mod.output("repriced.xml");
gen.parseFile("repriced.xml", false);
System.out.println("repriced.xml, same token: " + gen.getNav().toString(i));
}
}Output
238 tokens, 3858 bytes token 126 at 2031: 24.00 repriced.xml, same token: 26.00
Re-parsing the output puts 26.00 at the same token index: only bytes changed. The index costs roughly a third to a half of the file size, far less than a DOM, but there is no DOM or SAX API, no external entities, only XPath 1.0, a GPL licence and no release in nine years: an idea worth knowing rather than a dependency to add.
VTD XML
VTD-XML 7,086 (Virtual Token Descriptor for XML) is a group of cross-platform, non-extractive processing methods for XML. In traditional, extractive parsing, a lexical analyzer represents each token as a discrete string object copied out of the source. VTD-XML instead keeps the source document intact in memory and describes each token by its offset and length within that source - an approach its developers describe as "document-centric".Because tokens are recorded as compact numeric records rather than materialized as separate string/object instances, VTD-XML skips the object-oriented modeling step used by parsers such as DOM, eliminating the associated object-creation and garbage-collection costs. This lets navigation stay fast even on large documents, and supports incremental updates - modifying a document in place without a full re-parse. The tradeoffs are that the in-memory representation increases the effective document size by roughly 30% to 50%, and VTD-XML is not compatible with DOM, SAX, external DTD entities, or certain validation techniques.
Navigation in VTD-XML is performed through a cursor object (commonly named VTDNav in the Java and C/C# implementations), with an AutoPilot helper class used to run XPath-like expressions over the indexed document.
| Full name | Virtual Token Descriptor for XML |
| Processing model | Non-extractive; tokens recorded as offset/length pairs instead of extracted strings |
| Source document | Kept intact in memory alongside the token index |
| Memory overhead | Approximately 30% to 50% larger than the raw document size |
| Compatibility | Not compatible with DOM, SAX, external DTD entities, or certain validation techniques |
| Typical uses | Fast random access and incremental (in-place) updates without a full re-parse |
VTDGen vg = new VTDGen();
if (vg.parseFile("books.xml", false)) {
VTDNav vn = vg.getNav();
AutoPilot ap = new AutoPilot(vn);
ap.selectXPath("//book/@id");
int i;
while ((i = ap.evalXPath()) != -1) {
System.out.println("Book id: " + vn.toString(i));
}
}