In more detail
The file lives at a fixed address, such as https://example.com/robots.txt. It contains groups of rules: a User-agent line naming the crawler, followed by Disallow and Allow lines listing URL paths. It can also point to your XML sitemap. Google's introduction to robots.txt explains the syntax and how Google interprets it.
Key points from Google's documentation:
- Robots.txt is mainly for managing crawler traffic and avoiding crawling of unimportant or similar pages
- It isn't a way to keep a page out of Google. For that, use noindex, password protection or removal
- Blocking a page in robots.txt prevents Google from seeing a noindex on it
- Blocking CSS or JavaScript files can stop Google rendering pages correctly
Robots.txt is also how site owners control many AI crawlers. Google's own AI features in Search are governed by normal Googlebot rules, according to its AI features documentation. A separate product token, Google-Extended, controls use of content for some of Google's other AI systems. Other AI providers publish their own crawler names.
Mistakes in robots.txt can be severe. A single line such as Disallow: /, often left over from a staging site, blocks an entire site from crawling.
Why it matters when hiring an agency
Every audit should check robots.txt, and every site launch should include a check that the staging rules are gone. Ask an agency what it would block and why. On most small sites, very little needs blocking. On large sites, careful rules for filters and internal search can help crawl budget.
If you want to decide how AI crawlers use your content, ask the agency to explain the trade-offs for each crawler rather than blocking or allowing everything by default. Our Google Index Checker tests whether Googlebot may fetch a given URL under your robots.txt rules.