Every parser in Python and XML with lxml-Native XML Parser Libraries runs the same pipeline. Bytes are decoded into characters (the encoding comes from a byte-order mark, the XML declaration or the transport), a tokenizer cuts them into markup and text, a stack checks that the tokens nest, and the result is either handed to you as events or built into a tree, optionally checked against a schema on the way:
