xmlschema 476 4.3.2 (github.com/sissaschool/xmlschema (https://github.com/sissaschool/xmlschema 476 ), MIT) is a pure Python XSD 1.0 and 1.1 processor, used in Validating with an XSD to list a feed's errors. For data engineering its other half matters as much: it decodes a document into typed Python data driven by the schema, and encodes data back into XML, validating both ways:
import json
import xmlschema
schema = xmlschema.XMLSchema11("catalog11.xsd")
data = schema.to_dict("booknest-catalog.xml") # typed by the schema
b1 = data["book"][0]
print(type(b1["pages"]).__name__, type(b1["supply"]["price"]["$"]).__name__, b1["@id"])
rows = [{"sku": b["identifier"]["$"], "title": b["title"],
"price": float(b["supply"]["price"]["$"])} for b in data["book"]]
print(json.dumps(rows[1]))
b1["supply"]["price"]["$"] = -1 # encode back, with validation
elem, errors = schema.encode(data, path="catalog", validation="lax")
print(elem.tag, len(elem), "|", len(errors), "error:", errors[0].reason)Output
int Decimal b1
{"sku": "BN-0002", "title": "Patterns of the Deep Web", "price": 39.5}
catalog 7 | 1 error: value has to be greater than Decimal('0')Pages came back as int and prices as Decimal because the schema says so; attributes are @ keys and text content is $. Flat rows for Parquet 129 (JSON, Columnar and Binary Formats) are one comprehension away. With validation="lax", encoding returns errors instead of raising.