Dialects and Quoting

Dialects, Delimiters and Quoting

Real files vary in delimiter (semicolons where the comma is the decimal separator, tabs, pipes), quote character, escaping (doubled quotes or backslashes), quoting policy and line terminator. Python calls the combination a dialect (excel, excel-tab and unix are predefined, and csv.Sniffer guesses one); DuckDB 61,228 's sniff_csv detects dialect, header and column types together. Types are where readers disagree most:

DuckDB sniffs a European file; three readers type a ZIP code
printf 'id;title;price\r\n3;Salt and Saffron;24,00\r\n6;Gardens in Glass;21,30\r\n' > eu.csv
duckdb -line -c "SELECT Delimiter, HasHeader FROM sniff_csv('eu.csv')"
printf 'customer_id,zip\n1,02134\n' > zips.csv
duckdb -noheader -list -c "SELECT 'duckdb', zip, typeof(zip) FROM read_csv('zips.csv')"
python -c "import pandas as pd; print('pandas', pd.read_csv('zips.csv')['zip'][0])"
yq -p csv -o json -I 0 '.[0]' zips.csv
Output
Delimiter = ;
HasHeader = true
duckdb|02134|VARCHAR
pandas 2134
{"customer_id":1,"zip":2134}

DuckDB 1.5.5 kept ZIP code 02134 as text, while pandas 2.3.3 16,086 and yq 16,045 turned it into the integer 2134. Sniffers guess from samples, so in production declare the dialect and column types explicitly and agree them with the producer as part of a data contract.