Strings must escape ", \ and the control characters U+0000 to U+001F; anything else may appear literally or as a \uXXXX escape naming a UTF-16 code unit, so a character beyond the Basic Multilingual Plane needs a surrogate pair. Python escapes all non-ASCII by default: json.dumps({"note": "Café 📚"}) returns {"note": "Caf\u00e9 \ud83d\udcda"}, safe but two to three times larger for non-Latin text, so pass ensure_ascii=False and write UTF-8, which RFC 8259 requires between systems. The RFC admits its grammar allows a lone surrogate such as "\uDEAD", which encodes no character; Python parses it, then raises "surrogates not allowed" when the value reaches a UTF-8 file, far from its source. Implementations "MUST NOT add a byte order mark", yet PowerShell 5.1 55,538 's -Encoding UTF8 does, and json.loads fails with "Unexpected UTF-8 BOM"; decode with utf-8-sig. When embedding JSON in HTML, escape < as \u003c so no string can close a </script> tag.
MENU
Unicode and Escaping
Unicode, Escaping and Encoding Pitfalls