The Protocol That Ran the Internet for 28 Years Without Being a Standard
For 28 years, every major search engine crawler in the world followed a protocol that had never been ratified by any standards body. The robots.txt file, which tells web crawlers which pages they are allowed to access, operated on a voluntary system defined by a single informal document posted to a mailing list in 1994. It became foundational web infrastructure through adoption alone, not through formal specification. RFC 9309, which finally codified the Robots Exclusion Protocol, was published by the Internet Engineering Task Force in September 2022.
That 28-year gap between creation and standardization is not an anomaly. It reflects how much of the web's core infrastructure was built: one developer solving a concrete problem, other developers observing the solution and copying it, and the accumulated weight of adoption eventually making the convention as binding as any formal rule. Understanding what the informal era produced, and what the 2022 RFC finally resolved, explains why robots.txt configuration still trips up experienced developers.
The informal period is worth examining closely because it created inconsistencies that persist today, even among compliant, well-intentioned crawlers. Two search engines can both claim to support robots.txt and still interpret the same file differently on edge cases, because for most of the protocol's history there was no authoritative document to break the tie.
How a Single Mailing List Post Built Web Infrastructure
Martijn Koster created the original Robots Exclusion Protocol in February 1994, while working at Nexor, a UK networking company. His motivation was not abstract: automatic clients were overwhelming his website, making repeated requests that stressed the server. Koster posted a proposal to the www-talk mailing list, the primary communication channel for web-related technical discussions at the time, outlining a simple text file convention that crawlers should check before visiting any part of a site.
The proposal was deliberately minimal. A file named robots.txt at the domain root. A User-agent field identifying which crawler a rule applied to. A Disallow field listing paths that crawler should not visit. A wildcard asterisk in the User-agent field to match all crawlers. An empty Disallow value to signal full permission. The entire specification fit on a single page.
Adoption moved quickly. The major crawlers of 1994 implemented the convention within months of Koster's post. By June of that year it had become, as the RFC would later describe it, "a generally agreed upon de facto convention." Google, which launched in 1998, built robots.txt compliance into its crawler design from the beginning. When Googlebot encountered a robots.txt file, it followed the rules. When it did not find one, it crawled freely. The voluntary nature worked because crawler developers understood that non-compliance would result in the web community blacklisting their tools.
What the original document did not define became a list of edge cases that each crawler resolved independently. What should happen when different User-agent blocks gave conflicting instructions about the same URL? What did an asterisk within a path string mean? How long should a crawler wait before giving up on a temporarily unavailable robots.txt? How should it behave in the interim? Over the next 28 years, Google, Bing, Yahoo, and other search engines made their own decisions on each question, publishing documentation that often differed from what competitors were doing.
What RFC 9309 Resolved
Google submitted a draft to the IETF for formal standardization in 2019, and RFC 9309 was published in September 2022. Martijn Koster, the original author, was listed as a co-author of the RFC alongside engineers from Google, reflecting the continuity from the original 1994 proposal.
The RFC resolved several specific ambiguities. On rule conflicts, it defined that when a URL matches both an Allow and a Disallow directive, the longer, more specific rule wins, with ties broken in favor of Allow. This settled behavior that previously varied between crawlers. It clarified that URL paths in robots.txt are case-sensitive, so /Admin and /admin are treated as different paths, but that User-agent names are compared case-insensitively, so Googlebot and googlebot match the same rule.
Error handling received its first formal definition. When a server returns a 5xx error for robots.txt, the RFC specifies that crawlers should treat the entire site as disallowed until the file becomes available again. A 4xx error other than 429 (too many requests) means the crawler may proceed as if no robots.txt exists. A 429 rate-limiting response should cause the crawler to retry with reduced frequency. These were behaviors that different crawlers had handled differently, and the RFC gave them a canonical answer.
One area the RFC did not fully resolve is the Crawl-delay directive, which requests that crawlers wait a specified number of seconds between requests to a site. This directive was not in Koster's original 1994 document and is not part of RFC 9309. Google's documentation states explicitly that Googlebot does not honor Crawl-delay and manages its own crawl rate through internal mechanisms. Bing supports the directive. This means a robots.txt file that relies on Crawl-delay to protect server capacity provides that protection only for some crawlers, not all.
The Disallow Versus Noindex Confusion
One of the most consequential misunderstandings in robots.txt usage is the assumption that Disallow prevents a page from appearing in search results. It does not. Disallowing a path in robots.txt tells compliant crawlers not to visit that URL, but it does not prevent the URL from appearing in a search index by reference.
If a page that is disallowed in robots.txt has external links pointing to it from other pages that crawlers can visit, those links carry the URL to the index without the crawler ever reading the disallowed page's content. The URL may appear in search results with no title or description, described as a page that cannot be crawled. This is a documented behavior of Googlebot: it acknowledges URLs discovered through links even when robots.txt prevents it from visiting them.
The correct approach depends on the goal. To prevent crawling but allow indexing, use Disallow in robots.txt. To prevent indexing but allow crawling, add a noindex directive in the page's HTML head as a meta robots tag. To prevent both, allow crawling and use noindex, because a page that blocks crawlers cannot reliably receive the noindex signal the crawler needs to read in order to act on it.
Common Configuration Patterns
For most websites, a robots.txt file is short and focused. The standard structure allows all crawlers access everywhere by default and lists specific paths to disallow: admin interfaces, login pages, internal search result pages, duplicate pages generated by URL parameters, shopping cart and checkout flows, and private user content areas.
The Sitemap directive, added as a widely supported extension to the original protocol, allows robots.txt to reference XML sitemap files. Google, Bing, and most major crawlers support this directive and use it to discover pages they might not reach through link following alone. A robots.txt file for a large site typically ends with one or more Sitemap lines pointing to sitemap index files.
Wildcard patterns in the Disallow path use an asterisk to match variable URL segments. The pattern /search?* disallows all URLs beginning with /search? followed by any characters. The dollar sign at the end of a pattern anchors it to the end of the URL, so /page.pdf$ matches only URLs that end exactly with /page.pdf rather than any URL containing that string. Both wildcards are supported by Google and Bing but were only formally defined after the RFC process.
Conclusion
ToolHQ's robots.txt generator walks through these configuration cases and produces a correctly formatted file ready to deploy at your site root. The key value it provides is not in writing the basic User-agent and Disallow lines, which are simple enough to write by hand, but in handling the Sitemap directive, wildcard path syntax, and multi-crawler targeting in a format that all major crawlers will parse correctly according to RFC 9309.
The 28-year gap between informal convention and formal standard did not break the web. But it did leave a legacy of edge cases, crawler-specific behaviors, and common misconceptions that persist even now that the RFC exists. A correctly configured robots.txt file starts with understanding what the protocol actually specifies and where it still leaves room for interpretation.
Frequently Asked Questions
When was robots.txt officially standardized?
RFC 9309 was published by the IETF in September 2022, 28 years after Martijn Koster created the original informal Robots Exclusion Protocol in 1994.
Does disallowing a URL in robots.txt prevent it from appearing in Google?
No. Disallowing prevents Googlebot from crawling the page, but if other pages link to it, the URL may still appear in results as an unvisited link. Use a noindex tag to prevent search result appearances.
Does Google support the Crawl-delay directive in robots.txt?
No. Google explicitly states that Googlebot does not honor Crawl-delay and manages its own crawl rate. Bing and some other crawlers do support the directive.