A URL may contain only a limited set of ASCII characters, and some of them (/ ? # & =) give it structure. Any other byte is percent-encoded as % plus two hex digits, and non-ASCII text is first converted to UTF-8 bytes, so é becomes %C3%A9. Older tables that map é to %E9 show Windows-1252 bytes, which modern servers misread. Only A-Z a-z 0-9 - . _ ~ (RFC 3986 "unreserved") are safe unencoded everywhere.
| Char | Encoded | Char | Encoded | Char | Encoded |
|---|---|---|---|---|---|
| space | %20 (or + in forms) | & | %26 | / | %2F |
| " | %22 | + | %2B | : | %3A |
| # | %23 | , | %2C | = | %3D |
| % | %25 | ; | %3B | ? | %3F |
| @ | %40 | [ ] | %5B %5D | é / € | %C3%A9 / %E2%82%AC |
In JavaScript, encodeURIComponent() escapes everything except A-Z a-z 0-9 - _ . ! ~ * ' ( ), so use it for a single query value or path segment. encodeURI() leaves & and ? alone, which corrupts values containing them. Better still, let URL and URLSearchParams do it; they use form encoding, in which a space becomes +. Decode exactly once: decoding twice turns an attacker's %253C into <.
const q = 'café & crème/50% off?';
console.log(encodeURIComponent(q));
console.log(encodeURI(`https://example.com/search?q=${q}`)); // & splits the query: bug
const url = new URL('https://example.com/search');
url.searchParams.set('q', q);
console.log(url.href);caf%C3%A9%20%26%20cr%C3%A8me%2F50%25%20off%3F https://example.com/search?q=caf%C3%A9%20&%20cr%C3%A8me/50%25%20off? https://example.com/search?q=caf%C3%A9+%26+cr%C3%A8me%2F50%25+off%3F