SAX (Simple API for XML, a de facto standard from the XML-DEV mailing list, 1998) keeps no document: the parser pushes startElement, characters and endElement events into a handler, which keeps whatever state it needs, so memory stays flat. A classic bug is assuming one characters() call per text node; parsers split text at buffer boundaries, so accumulate it:
import java.io.File;
import javax.xml.parsers.SAXParserFactory;
import org.xml.sax.Attributes;
import org.xml.sax.helpers.DefaultHandler;
public class Sax extends DefaultHandler { // the parser calls these methods
private final StringBuilder text = new StringBuilder();
private double total;
@Override public void startElement(String uri, String local, String name, Attributes a) {
text.setLength(0);
}
@Override public void characters(char[] ch, int start, int length) {
text.append(ch, start, length); // may arrive in several pieces
}
@Override public void endElement(String uri, String local, String name) {
if (name.equals("price")) total += Double.parseDouble(text.toString());
}
public static void main(String[] args) throws Exception {
Sax h = new Sax();
SAXParserFactory.newInstance().newSAXParser().parse(new File(args[0]), h);
System.out.printf("SAX: total %.2f%n", h.total);
}
}SAX
SAX (Simple API for XML) parsers use streaming: rather than building a tree for the whole document, they operate on each piece of the document sequentially, in a single pass, and feed the result directly to the application. Where DOM operates on the document as a whole, SAX is state-independent - it does not build or retain a representation of the document.SAX is event-driven and push-based: the parser fires callback events, such as start-of-element, character data, and end-of-element, as it encounters them, and the application's event handlers respond to each event as it arrives. Because no document tree is retained, the memory required by a SAX parser is minimal, and processing is generally faster than DOM for large documents. The tradeoff is that the application, not the parser, is responsible for tracking any state it needs across events, such as which element is currently open.
SAX originated as a de facto standard for Java, developed collaboratively by the XML-DEV mailing list rather than by a formal standards body, and has since been ported to many other languages. It is well suited to processing large documents in a single pass, or to extracting a small amount of information without needing random access.
| Full name | Simple API for XML |
| Processing model | Event-driven streaming, push-based |
| State | Stateless at the parser level; the application tracks its own state |
| Memory use | Minimal; the document is not retained in memory |
| Typical uses | Single-pass processing of large documents; extracting selected data without full-document access |
| Related | Implemented by Xerces and other toolkits; contrasts with pull-based StAX |
import xml.sax
class BookHandler(xml.sax.ContentHandler):
def startElement(self, name, attrs):
if name == "book":
print("Book id:", attrs.get("id"))
xml.sax.parse("books.xml", BookHandler())