Safe Defaults

Safe Defaults for Untrusted Input

Most parsers ship limits; the work is knowing which are on by default. Four hostile inputs:

guards.py: limits that stop hostile input earlyPython
import json, msgpack
tests = {"5,000-digit number": lambda: json.loads("9" * 5000),
         "100,000 nested arrays": lambda: json.loads("[" * 100_000 + "]" * 100_000),
         "msgpack 4 GB claim": lambda: msgpack.unpackb(bytes.fromhex("dbffffffff") + b"x"),
         "msgpack 2 MB string": lambda: msgpack.unpackb(msgpack.packb("x" * 2**21),
                                                        max_str_len=2**20)}
for name, test in tests.items():
    try:
        print(f"{name:22} accepted: {test()}")
    except Exception as e:
        print(f"{name:22} {type(e).__name__}: {str(e)[:55]}")
Output
5,000-digit number     ValueError: Exceeds the limit (4300 digits) for integer string conv
100,000 nested arrays  RecursionError: Stack overflow (used 8160 kB) while decoding a JSON arr
msgpack 4 GB claim     ValueError: Unpack failed: incomplete input
msgpack 2 MB string    ValueError: 2097152 exceeds max_str_len(1048576)

Three limits were built in: Python's 4,300-digit cap (added for CVE-2020-10735), its stack guard, and msgpack's check of a declared length against the bytes present. max_str_len had to be chosen. For every ingestion endpoint, cap sizes before parsing, reject what the spec forbids (NaN, duplicate keys), validate against a schema before building objects, and parse in a memory-limited worker, so a bomb kills a process, not the pipeline.