A character set maps characters to bytes; a collation compares and sorts them. The default is utf8mb4 (full UTF-8, 1 to 4 bytes per character) with utf8mb4_0900_ai_ci: Unicode 9.0 rules, accent- and case-insensitive. The deprecated utf8mb3, which the bare name utf8 still means, stops at 3 bytes and cannot store emoji.
SET @s = CONCAT('Café ', _utf8mb4 0xF09F9098); -- 0xF09F9098 is U+1F418, an emoji
CREATE TABLE cs (old VARCHAR(20) CHARACTER SET utf8, new VARCHAR(20));
INSERT INTO cs (old) VALUES (@s);
SELECT CHAR_LENGTH(@s) AS chars, LENGTH(@s) AS bytes, 'resume' = 'Résumé' AS ai_ci,
'resume' = 'Résumé' COLLATE utf8mb4_0900_as_ci AS as_ci,
'Resume' = 'resume' COLLATE utf8mb4_0900_as_cs AS as_cs,
'ß' = 'ss' AS ss_0900, 'ß' = 'ss' COLLATE utf8mb4_general_ci AS ss_general;Warning (Code 3719): 'utf8' is currently an alias for the character set UTF8MB3, but will be an alias for UTF8MB4 in a future release. Please consider using UTF8MB4 in order to be unambiguous. ERROR 1366 (HY000): Incorrect string value: '\xF0\x9F\x90\x98' for column 'old' at row 1 +-------+-------+-------+-------+-------+---------+------------+ | chars | bytes | ai_ci | as_ci | as_cs | ss_0900 | ss_general | +-------+-------+-------+-------+-------+---------+------------+ | 6 | 10 | 1 | 0 | 0 | 1 | 0 | +-------+-------+-------+-------+-------+---------+------------+
Six characters take ten bytes (é two, the emoji four). The default collation finds "resume" equal to "Résumé", right for search but wrong for a case-sensitive coupon code, which needs utf8mb4_0900_as_cs or utf8mb4_bin; the legacy utf8mb4_general_ci even disagrees on ß. Use one collation throughout, with charset=utf8mb4 in the PHP DSN (Character Sets): joining a utf8mb4_0900_ai_ci column to a utf8mb4_general_ci one fails with ERROR 1267, "Illegal mix of collations".