Most parsers ship limits; the work is knowing which are on by default. Four hostile inputs:
import json, msgpack
tests = {"5,000-digit number": lambda: json.loads("9" * 5000),
"100,000 nested arrays": lambda: json.loads("[" * 100_000 + "]" * 100_000),
"msgpack 4 GB claim": lambda: msgpack.unpackb(bytes.fromhex("dbffffffff") + b"x"),
"msgpack 2 MB string": lambda: msgpack.unpackb(msgpack.packb("x" * 2**21),
max_str_len=2**20)}
for name, test in tests.items():
try:
print(f"{name:22} accepted: {test()}")
except Exception as e:
print(f"{name:22} {type(e).__name__}: {str(e)[:55]}")Output
5,000-digit number ValueError: Exceeds the limit (4300 digits) for integer string conv 100,000 nested arrays RecursionError: Stack overflow (used 8160 kB) while decoding a JSON arr msgpack 4 GB claim ValueError: Unpack failed: incomplete input msgpack 2 MB string ValueError: 2097152 exceeds max_str_len(1048576)
Three limits were built in: Python's 4,300-digit cap (added for CVE-2020-10735), its stack guard, and msgpack's check of a declared length against the bytes present. max_str_len had to be chosen. For every ingestion endpoint, cap sizes before parsing, reject what the spec forbids (NaN, duplicate keys), validate against a schema before building objects, and parse in a memory-limited worker, so a bomb kills a process, not the pipeline.