
Adam Wozniak
My starting observation is that a URL is not one string. It is a set of components with separate rules, and the parser on the other end may split it differently than you expect. That mismatch, not a mistake in the encoding itself, is where most URL defects occur. Encoding has a defined safe list. Unreserved characters survive untouched; everything else is percent-encoded, and the encoded form is case-insensitive for the hex digits but case-sensitive for what it represents. The two rules people get wrong are that a space encodes to a plus sign in form data and to a percent sequence in a path, and that encoding an already-encoded string double-encodes it. The query string is where the structure disappears. Parameters are separated by an ampersand, keys and values by an equals sign, and neither the order nor the repetition of a key carries any defined meaning. Two parsers will disagree about an unencoded ampersand inside a value, and neither is wrong. The fragment is the part people most often misunderstand. It is never sent to the server, so anything placed after it cannot reach the backend at all, and putting a token there means it does not arrive. Internationalised domains become punycode, which is worth recognising in a log line. A name displayed as an accent and a name displayed as its encoded form are the same host, and a filter comparing the two as strings will not agree. I also cover the boring limits: maximum length in browsers and servers, characters that are legal in a path and rejected by a proxy in front of it, and how a single trailing slash change produces two URLs the server treats as different pages.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →