A CSV file does not declare its encoding. Excel's "CSV UTF-8" writes a byte order mark (EF BB BF), and older Windows exports use code page 1252. Both break a reader that assumes plain UTF-8:
import csv
text = "id,title,zip\n1,Café Stories,02134\n"
open("excel.csv", "w", encoding="utf-8-sig", newline="").write(text) # what Excel saves
open("legacy.csv", "w", encoding="cp1252", newline="").write(text) # old Windows export
for enc in ("utf-8", "utf-8-sig"):
with open("excel.csv", encoding=enc, newline="") as f:
print(f"{enc:9}", next(csv.DictReader(f)))
print("UTF-8 read as cp1252:", open("excel.csv", encoding="cp1252").read().split()[1])
try:
open("legacy.csv", encoding="utf-8").read()
except UnicodeDecodeError as err:
print("cp1252 read as UTF-8:", err.reason, "at byte", err.start)Output
utf-8 {'\ufeffid': '1', 'title': 'Café Stories', 'zip': '02134'}
utf-8-sig {'id': '1', 'title': 'Café Stories', 'zip': '02134'}
UTF-8 read as cp1252: 1,Café
cp1252 read as UTF-8: invalid continuation byte at byte 18With plain utf-8 the BOM joins the first header, so row["id"] raises KeyError; UTF-8 read as cp1252 is silent mojibake. Other classic breakages: empty string versus null (to_csv.py writes a missing coupon as an empty field), locale-dependent dates and decimals, formula injection when a field starts with =, ragged rows and inserted columns. Decode with utf-8-sig, fail on bad bytes, declare types, count fields, and keep CSV for human handoffs, not for pipelines between machines.