xmlschema Validation

Declarative Validation with xmlschema

xmlschema 476 4.3.2 (github.com/sissaschool/xmlschema (https://github.com/sissaschool/xmlschema 476 ), MIT) is a pure Python XSD 1.0 and 1.1 processor, used in Validating with an XSD to list a feed's errors. For data engineering its other half matters as much: it decodes a document into typed Python data driven by the schema, and encodes data back into XML, validating both ways:

decode.py: schema-typed dictionaries, flat rows and validated encodingPython
import json
import xmlschema
schema = xmlschema.XMLSchema11("catalog11.xsd")
data = schema.to_dict("booknest-catalog.xml")              # typed by the schema
b1 = data["book"][0]
print(type(b1["pages"]).__name__, type(b1["supply"]["price"]["$"]).__name__, b1["@id"])
rows = [{"sku": b["identifier"]["$"], "title": b["title"],
         "price": float(b["supply"]["price"]["$"])} for b in data["book"]]
print(json.dumps(rows[1]))
b1["supply"]["price"]["$"] = -1                             # encode back, with validation
elem, errors = schema.encode(data, path="catalog", validation="lax")
print(elem.tag, len(elem), "|", len(errors), "error:", errors[0].reason)
Output
int Decimal b1
{"sku": "BN-0002", "title": "Patterns of the Deep Web", "price": 39.5}
catalog 7 | 1 error: value has to be greater than Decimal('0')

Pages came back as int and prices as Decimal because the schema says so; attributes are @ keys and text content is $. Flat rows for Parquet 129 (JSON, Columnar and Binary Formats) are one comprehension away. With validation="lax", encoding returns errors instead of raising.