RFC 4180 (October 2005) is an Informational RFC that documented existing practice and registered text/csv. Its seven rules: records end in CRLF; the final line break is optional; an optional header looks like a record; fields are comma-separated, records should have equal field counts, and spaces are data; fields may be quoted; fields containing CRLF, quotes or commas should be quoted; and a quote inside a quoted field is doubled. It says nothing about types, nulls or encoding beyond a charset parameter.
import csv
rows = [["book_id", "reviewer", "comment"],
[5, "Rosa Silva", 'Twisty, clever, "unputdownable"'],
[3, "Ito, Jia", "Line one\nline two"]]
with open("reviews.csv", "w", newline="", encoding="utf-8") as f:
csv.writer(f).writerows(rows) # excel dialect: CRLF, quote only when needed
with open("reviews.csv", newline="", encoding="utf-8") as f:
print("csv.reader fields per record:", [len(r) for r in csv.reader(f)])
with open("reviews.csv", encoding="utf-8") as f:
print("split(',') fields per line: ", [len(line.split(",")) for line in f])python csv_rfc.py && cat -A reviews.csvcsv.reader fields per record: [3, 3, 3]
split(',') fields per line: [3, 5, 4, 1]
book_id,reviewer,comment^M$
5,Rosa Silva,"Twisty, clever, ""unputdownable"""^M$
3,"Ito, Jia","Line one$
line two"^M$^M$ marks a CRLF; the embedded line break stays a bare LF inside quotes. A parser sees three records of three fields, split(",") four lines of 3, 5, 4 and 1. Never parse CSV with split, cut or a regular expression.