
Ben Halvorsen
I start by explaining why every CSV problem exists, which is that the format has no way to distinguish a delimiter inside a value from a delimiter between values, so it invents rules and relies on everybody following them. Quoting is those rules. A value containing the delimiter, a quote or a newline has to be quoted, and a quote inside a quoted value has to be doubled. A generator that does one and not the other produces a file that opens in one tool and breaks in the next. A newline inside a quoted value is the case worth showing in full, because it splits one row into two on any reader that does not implement the rule, and a spreadsheet application is exactly such a reader on a bad day. Encodings decide whether the file opens at all. A byte order mark at the start is read as data by some tools and as a marker by others, and a file in one encoding read as another produces a first line containing several garbage characters. That first line is usually the header, so the garbage lands in the first column name. Ragged rows are legal. A short row fills the missing columns with nothing and a long row adds unnamed columns, and both behaviours differ between readers. I finish on dialect differences, because the separator, the quote character, the escape and the null representation are all configurable and all frequently vary between producers. I put every example in the page as text so the delimiters are visible, because a CSV problem is invisible in a spreadsheet and obvious the moment you look at the bytes.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →