
Nina Volkova
The subject is small and the failure modes are specific, so I go through them one at a time. Turkish dotted and dotless i are the standard example, and a naive conversion produces a different letter rather than a different form. Title case is where the disagreement lives. Converting to title case requires deciding what counts as a word and which words get capitalised. A dictionary-based approach knows that prepositions are usually lowercased unless they lead or follow, and a rule-based approach cannot. Headlines follow a different convention from prose, and a style guide that does not say which one applies leaves it to the tool. Sentence case is simpler and is the convention most web copy uses. One complication: an abbreviation or a proper noun needs restoring afterwards, since capitalising the first letter of every sentence breaks anything that was already an acronym. Toggle case is the one everybody uses by accident. It swaps the case of every letter, which is useful for shouting emphasis and is the wrong tool for anything else. It also breaks text that contains a mixture deliberately, such as an identifier. I finish with the Unicode point rather than the byte, because converting bytes gives you an error on any non-ASCII character and converting code points gives you the character. There is also the question of what happens to characters with no case at all, and to scripts where one character form is used for two purposes, which is where a conversion that looks correct produces something wrong. Script-specific rules exist for this reason, and a tool that applies one language's rules to another will get it wrong quietly. Where a language's case rules do not reduce to a simple mapping, a general-purpose converter will still produce a plausible result. That is the case I would check by hand.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →