buf.toString() needs to know how the bytes encode characters. Node's built-in set is utf8 (the default), utf16le (aliased ucs2), latin1 (misleadingly aliased binary), ascii, and the binary-to-text codings hex, base64 and base64url. Anything else — windows-1252, shift_jis, gb18030 — needs TextDecoder and its full WHATWG encoding set.
The trap is multi-byte characters split across chunk boundaries: a stream hands you arbitrary byte counts, so é or 世 may straddle two chunks.
import { StringDecoder } from 'node:string_decoder';
const bytes = Buffer.from('héllo 世界', 'utf8'); // 13 bytes
const [a, b] = [bytes.subarray(0, 8), bytes.subarray(8)];
console.log('naive:', a.toString('utf8') + b.toString('utf8'));
const decoder = new StringDecoder('utf8');
console.log('decoder:', decoder.write(a) + decoder.write(b) + decoder.end());
const td = new TextDecoder('utf-8', { fatal: true });
console.log('stream:', td.decode(a, { stream: true }) + td.decode(b));naive: héllo ���界 decoder: héllo 世界 stream: héllo 世界
Both hold the incomplete tail until the next chunk completes it. You rarely call either by hand: createReadStream(path, { encoding: 'utf8' }), setEncoding('utf8') and TextDecoderStream use one internally.
Pass fatal: true when the bytes must be valid: without it a decoder substitutes U+FFFD, turning corrupt input into plausible-looking text. isUtf8(buf) and isAscii(buf) from node:buffer check cheaply before you decode an upload; latin1 never fails, which is why it is the wrong default for user content.