
Sunita Rao
Every cleaning decision needs to be stated as a decision, because the same file cleaned by two people comes out different and the difference is nearly invisible. The second kind needs a definition of similarity and a rule for which copy survives, and both choices change the result of everything downstream. Removing the first rather than the last assumes an ordering that may not exist. Column names are inconsistent more often than values are. Mixed case, and two spellings of the separator within one file, and abbreviations for the same concept across files imported at different times are the usual causes, and a mapping applied once at load time fixes what nobody fixes. Missing values have many spellings: empty, a single space, a null token, a literal word for unknown, and a sentinel like minus one or zero. They mean different things and merging them loses the distinction, so the mapping should keep them apart until you know which are the same. Types are inferred by whatever reads the file, and a column with one odd value becomes text. That is a parsing decision rather than a data one, and it is worth declaring the expected types instead of accepting the inference. I end on validation, because a cleaning step with no check afterwards is a cleaning step whose effect nobody can measure. The order of the steps matters too, because deduplicating before trimming means comparing rows that differ only in whitespace, and trimming before means the comparison is on the data you meant.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →