
Dorin Ionescu
I start from the failures that reach production having passed every local check, because that gap is where XML spends most of its time. The encoding declaration is the usual culprit. XML processors detect the character set from a byte order mark, from the declaration, or from the HTTP content type header, resolving conflicts in a specified order. A document saved as UTF-8 and declared as ISO-8859-1 parses without complaint and returns mojibake. Declaring nothing is legal, and UTF-8 is the default, which is another reason the declaration is worth writing. Namespaces are the other half. A prefix is bound by a namespace declaration and carries no meaning without it, so two documents can use the same prefix for different namespaces and the same namespace under different prefixes. Comparing elements by prefix rather than by namespace URI is a recurring bug. Entities divide into the five built-ins, character references written as an escape and a number, and everything else. A document type declaration can define further entities, and expanding them is how entity expansion attacks work. Parsers should be configured to forbid external entity resolution outright. CDATA sections exist because a parser treats angle brackets as markup everywhere, which makes ordinary prose containing markup awkward to include. A CDATA section suspends that. It does not nest, and it does not escape, so text containing the closing sequence still breaks. Validation is where I draw a firm line. Well-formedness is a parse result. Validity against a schema is a separate, optional step, and most production parsers never perform it.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →