A decision about training data can affect a different service from a decision about search visibility. A crawler name, a page directive and a Search Console setting can each govern a different part of the process.
Before changing access, write down the outcome you want: discoverable product pages, eligibility for AI search answers, a limit on displayed excerpts, exclusion from specified model training, or a private page that nobody outside the authorized audience should reach. Choose controls for that outcome and verify them at the deployed URL.
This reference covers the public controls documented by Google, Microsoft and OpenAI. It does not replace a provider contract or establish what an unrelated crawler will do.
The control matrix
| Provider and control | What the documentation says it governs | Operational distinction | Primary source |
|---|---|---|---|
| Googlebot in robots.txt | Crawling for Google Search and associated search features | Blocking the search crawler affects the foundation that search experiences use | Google crawler list |
| Google-Extended in robots.txt | Specified Gemini training and grounding uses | A control token without its own HTTP user-agent string; it does not determine Google Search inclusion or ranking | Google crawler list |
| Google Search generative AI setting | Inclusion in covered Search generative AI features | Inspect the effective property setting, including inheritance from a parent property | Search Console control |
| Google noindex and snippet directives | Search inclusion and the content available for previews or direct AI input, according to the directive | Google must be able to fetch the page to read its directives | Google page controls |
| Bingbot access and NOINDEX | Crawling; separately, inclusion in Bing search, Copilot and grounding results | A crawl restriction and an indexing instruction do different jobs | Bing current guidelines |
| Bing NOARCHIVE and NOCACHE | Copilot and grounding use | NOARCHIVE prevents the documented use; NOCACHE limits it to URL, title and snippet | Bing current guidelines |
| OpenAI OAI-SearchBot | ChatGPT search crawling | Search access is separate from the GPTBot training preference | OpenAI crawlers |
| OpenAI GPTBot | Content that may be used to train foundation models | Its robots.txt setting is independent of OAI-SearchBot | OpenAI crawlers |
| OpenAI ChatGPT-User | Certain visits initiated by a user | It is not the automatic search crawler; robots.txt may not apply to those user-driven actions | OpenAI crawlers |
Read the linked documentation before configuring a live site. This table summarizes scope; it is not a universal allowlist.
Google Search and Gemini require separate decisions
A request to “block Google AI” is too broad to implement safely. Ask which use the owner means. Google-Extended includes grounding in the specified Gemini products as well as training. Treating it as a training-only switch can produce an outcome the owner did not intend.
The Search generative AI control adds a separate property-level decision. A child property can inherit its setting, so checking only a subdomain's robots.txt leaves a gap. Record the effective setting and the property it comes from. Control and inheritance
If you want a page to remain in ordinary search while limiting excerpts, inspect the relevant page controls. Google's nosnippet rule also prevents content from being used as direct input for AI Overviews and AI Mode; max-snippet limits that input. Those effects deserve review before a site-wide change. Snippet rules
Bing directives can affect AI answers
Bing's current webmaster guidelines describe NOARCHIVE and NOCACHE as grounding controls. Do not assume they only concern a visible cached-page link. Review inherited CMS, paywall and security defaults before concluding that a page is available for the intended Copilot experience.
Microsoft's September 2023 announcement also described prospective training restrictions for content in its Bing index: NOARCHIVE excluded that content, while NOCACHE restricted use to the URL, title and snippet. The announcement uses the older Bing Chat name. Its stated scope should not be expanded into a claim about every Microsoft product or contract. Dated Microsoft announcement
A policy-sensitive training decision should be checked against the relevant current terms and support guidance. The crawler matrix alone cannot settle it.
OpenAI search access does not require a matching training choice
An owner can make different choices for OAI-SearchBot and GPTBot. OpenAI explicitly documents their independence. Its crawler page also publishes IP information for verification; a request carrying a familiar user-agent name is not sufficient evidence of identity. OpenAI controls and crawler verification
Keep the difference between an automatic crawler and a user-triggered visit in the access plan. If a document must be private, protect it with authentication and authorization rather than relying on a voluntary crawling preference.
Verify the whole delivery path
A robots file is one input. Use the deployed URL checklist for response, discovery, source text and payload checks. Keep the provider-policy decisions in a record like this:
| Check | Evidence to retain | Failure to investigate |
|---|---|---|
| Intended policy | Owner-approved service, path and purpose | “Allow AI” without specifying a provider or use |
| robots.txt | Deployed content, response status and applicable group | Wrong host, stale file, accidental broad restriction |
| Page directives | Effective meta and HTTP-header rules | A template adds noindex or a restrictive snippet rule |
| Provider controls | Effective property or account setting | A parent setting overrides the expected choice |
| Verified access | Relevant logs or official inspection evidence | A test user-agent was mistaken for a real crawler |
A successful local fetch proves that the test client received a response. It does not prove what a real crawler received from another network. Conversely, one tool failing to open a URL does not establish that a provider has blocked or deindexed it.
For Google, a 404 robots.txt response is treated as no crawl restrictions. Do not describe a missing file as a universal indexing block. Google robots.txt handling
Make changes reversible and reviewable
Test the proposed change on the affected scope, record the previous state, and agree how to restore it. Check both the policy file and representative pages after deployment. Use provider verification methods for genuine crawler traffic and review the effect after the provider has had an opportunity to process the change.
Keep access testing separate from the next question: whether the page is useful enough, relevant enough and correctly represented when an answer is generated. The product-data readiness checklist covers the factual record. The audit methodology covers the answer evidence.