UTF-8 stores each Unicode code point in one to four bytes, and one visible character, a grapheme cluster, can be several code points: a letter plus a combining accent, a flag, or a skin-toned emoji. Core functions count bytes, mb_* functions code points, and intl's grapheme_* functions what users see.

<?php
$s = "Cafe\u{0301} 日本"; // e + combining acute accent
echo strlen($s), ' bytes, ', mb_strlen($s), ' code points, ',
grapheme_strlen($s), " graphemes\n";
var_dump(mb_check_encoding(substr($s, 0, 5)), mb_substr($s, 0, 4), grapheme_substr($s, 0, 4));
echo json_encode(grapheme_str_split($s), JSON_UNESCAPED_UNICODE), ' ',
mb_str_pad('Zoë', 5, '.'), '|', str_pad('Zoë', 5, '.'), "|\n";
$legacy = "caf\xE9"; // Windows-1252 bytes from an old CSV
echo mb_convert_encoding($legacy, 'UTF-8', 'Windows-1252'), ' ', mb_scrub($legacy), ' ',
levenshtein('noël', 'noel'), ' ', grapheme_levenshtein('noël', 'noel'), ' ',
var_export(Normalizer::normalize("e\u{0301}") === 'é', true), "\n";13 bytes, 8 code points, 7 graphemes bool(false) string(4) "Cafe" string(6) "Café" ["C","a","f","é"," ","日","本"] Zoë..|Zoë.| café caf? 2 1 true
Cutting at byte 5 left invalid UTF-8, and four code points lost the accent; only grapheme_substr() kept "Café". Use mb_* for limits such as a VARCHAR(50) column (MySQL) and grapheme_* for text people read. PHP 8.3 added mb_str_pad(), 8.4 grapheme_str_split(), and 8.5 grapheme_levenshtein(). Validate with mb_check_encoding(), convert legacy bytes with mb_convert_encoding() (utf8_encode() is deprecated since 8.2), and normalize before comparing, since é can be one code point or two.