Multibyte Strings

Bytes, Characters and Multibyte Strings

UTF-8 stores each Unicode code point in one to four bytes, and one visible character, a grapheme cluster, can be several code points: a letter plus a combining accent, a flag, or a skin-toned emoji. Core functions count bytes, mb_* functions code points, and intl's grapheme_* functions what users see.

One string counted as graphemes, code points and UTF-8 bytes
One string counted as graphemes, code points and UTF-8 bytes
Three ways to count, cut and convert text (multibyte.php)PHP
<?php
$s = "Cafe\u{0301} 日本";                               // e + combining acute accent
echo strlen($s), ' bytes, ', mb_strlen($s), ' code points, ',
    grapheme_strlen($s), " graphemes\n";
var_dump(mb_check_encoding(substr($s, 0, 5)), mb_substr($s, 0, 4), grapheme_substr($s, 0, 4));
echo json_encode(grapheme_str_split($s), JSON_UNESCAPED_UNICODE), ' ',
    mb_str_pad('Zoë', 5, '.'), '|', str_pad('Zoë', 5, '.'), "|\n";
$legacy = "caf\xE9";                                    // Windows-1252 bytes from an old CSV
echo mb_convert_encoding($legacy, 'UTF-8', 'Windows-1252'), ' ', mb_scrub($legacy), ' ',
    levenshtein('noël', 'noel'), ' ', grapheme_levenshtein('noël', 'noel'), ' ',
    var_export(Normalizer::normalize("e\u{0301}") === 'é', true), "\n";
Output
13 bytes, 8 code points, 7 graphemes
bool(false)
string(4) "Cafe"
string(6) "Café"
["C","a","f","é"," ","日","本"] Zoë..|Zoë.|
café caf? 2 1 true

Cutting at byte 5 left invalid UTF-8, and four code points lost the accent; only grapheme_substr() kept "Café". Use mb_* for limits such as a VARCHAR(50) column (MySQL) and grapheme_* for text people read. PHP 8.3 added mb_str_pad(), 8.4 grapheme_str_split(), and 8.5 grapheme_levenshtein(). Validate with mb_check_encoding(), convert legacy bytes with mb_convert_encoding() (utf8_encode() is deprecated since 8.2), and normalize before comparing, since é can be one code point or two.