
Marcus Oyelaran
My first move on any indexing problem is to check what the server sends back, because the setting in the content system and the response the crawler received are frequently different things. There are two ways to say do not index, and both have to be in agreement. A robots meta tag in the page, and a header carrying the same instruction. The header takes precedence, so a page with a noindex tag and a missing header from a misconfigured rule will still be indexed, which is the single most common cause of a page that refuses to leave a result page. A page excluded through the robots file cannot be crawled, so its meta tag is never read, so it cannot be indexed that way either. The file and the tag solve different problems: the file prevents crawling, the tag prevents indexing while allowing the crawl. Removing a page from a sitemap does not remove it from an index. The sitemap is for discovery. If a page needs to leave an index it needs the instruction, and the instruction takes a crawl to be seen. Deletion is the slowest path, since removal depends on a crawler revisiting the address, which is an argument for keeping low-value pages rather than deleting them and hoping. I also cover the case that looks like an indexing bug and is not: a page indexed at a different address than the one requested, which is a canonical problem. Each of these is settled by the response the crawler received, so I show that response rather than the configuration that produced it, since the two diverge more often than the tooling suggests.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →