
Nikhil Shetty
I read these pages from the crawler's side, because the directive that matters is the one the crawler read rather than the one that was intended. The robots file is a set of path rules with user-agent groups. The matching is prefix-based, so a rule for one directory covers everything below it, and a rule for a path that appears twice in the file is resolved by the least specific one winning. That last rule surprises people who add an exception below a broad rule and find it ignored. The file does two things and conflating them causes most of the trouble. It controls crawling and it controls what may appear in a result. A page disallowed from crawling cannot have its robots tag read, so it cannot be indexed either, which means a disallowed page is invisible rather than merely excluded. The file should hold paths you want crawled but not indexed, with the tag on each page doing that work. Sitemaps declare URLs for discovery and optionally carry a last-modified date. A date that is wrong on every URL is worse than none, since a crawler uses it to decide what to re-fetch and a constant date means never re-fetching. Crawl budget is the framing I use for a large site. Internal links are the control: depth from the home page determines how many requests it takes to reach a page, and anything buried is reached rarely. Blocking a search engine's crawler blocks your own diagnostics, so it is worth knowing what it can tell you about a site before you exclude it.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →