
Hannah Vogel
My work is unglamorous and it decides everything downstream. A model given a ragged file does not return a better answer because it is a good model. Format choice comes first. Line-delimited JSON parses incrementally and survives a truncated record. CSV is what non-specialists can edit, and it breaks on values containing the delimiter, on newlines inside fields and on inconsistent quoting. Parquet and similar columnar formats carry types and compress well. Choosing a format by what the producer emits rather than what the consumer prefers produces a pipeline somebody maintains by hand. The structural problems I see most are ragged header rows, inconsistent date formats, numbers stored as strings with thousands separators, and nulls written as several different tokens. Each is cheap to fix once and expensive to fix downstream, and each is invisible when you only look at the first few rows. Deduplication needs a decision, not a function. Exact duplicates can go. Near duplicates need a similarity threshold and a policy for which one survives, and the choice affects your results more than most people expect. Identifiers and personal data deserve an explicit pass. Anything that identifies a person should be removed or replaced before a dataset goes anywhere, and pseudonymisation done at export is far easier than untangling it afterwards. Finally, the split. A dataset where evaluation examples appear in training will produce numbers that mean nothing, and that mistake is easy to make because nothing about the code announces it. One more thing worth checking before you trust a file: whether the encoding is what the extension claims. A UTF-8 file renamed from something else parses and produces replacement characters rather than an error, and those are easy to mistake for empty values.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →