Skip to main content
Technical readiness · OpenAI Crawlers

AI Search Crawler Access: Why OAI-SearchBot and Training Crawlers Need Separate Policies

A site can permit search discovery while expressing a different training preference. robots.txt is only one layer: CDN and WAF policies, origin verification, responses and logs determine actual access.

Key Takeaway

OAI-SearchBot supports discovery in ChatGPT search; GPTBot crawls content that may be used to improve and train foundation models; ChatGPT-User handles user-triggered visits rather than search indexing. These controls are independent, so a site may allow OAI-SearchBot while disallowing GPTBot. Verify genuine requests against official IP information and check CDN, WAF, response and log behavior without bypassing application security.

What's the official change?

OpenAI's cited documentation separates crawler purposes. Disallowing OAI-SearchBot opts page content out of search answers, though a navigational link may still appear. Sites seeking discovery should permit the crawler and avoid blocking verified requests from published official ranges.

GPTBot serves a different purpose: crawling content that may help improve foundation models. Its opt-out signal concerns potential training and is independent of search preferences. Where both purposes are permitted, OpenAI may reuse a crawl to avoid duplicate requests; administrators can still set separate policies. An opt-out is not a claim about removing all previously acquired data.

ChatGPT-User handles specific visits triggered by users of ChatGPT or custom GPTs. OpenAI notes that robots.txt may not apply to these user-directed requests and that this agent does not determine search inclusion. A generic 'ChatGPT bot' rule obscures these distinctions.

OAI-SearchBot ChatGPT search discovery and citations
GPTBot Potential model-training use
ChatGPT-User User-triggered page visits

Implications for Brands Expanding Internationally

Search goals and training preferences can be decided independently. A brand may allow OAI-SearchBot while deciding GPTBot access with its legal, intellectual-property and data owners. Blanket allow or block policies can unnecessarily couple different objectives.

An Allow rule does not prove successful retrieval. CDN bot controls, WAF rules, IP reputation, rate limits, geography, TLS and origin authentication can still return 403, 429, 5xx, challenge pages or empty content. Verify actual responses and server logs rather than only the robots file.

Review public-page scope before permitting discovery. Remove private customer data, unapproved specifications, internal test records and obsolete commitments. Use crawl-readable noindex for public pages that should not appear in results, and genuine authentication or authorization for confidential resources. A robots block alone does not prevent a URL or title appearing as a link.

Zhihe Growth's Assessment

AI bot access requires more than a robots rule. Audit four layers: policy for search, training and user-directed access; network checks covering official IP ranges, DNS, CDN and WAF settings; page status codes, canonicals, sitemaps, visible content, resource loading and noindex; and logs showing when the intended User-Agent reached which URLs, what responses it received and whether failures recurred.

Zhihe Growth recommends enabling intended search access only with the site's approved policy and obtaining explicit owner approval for training preferences. Do not block OAI-SearchBot merely because GPTBot is unwanted, or open training access for speculative GEO gains. Retain policy versions, dates, owners, approvals, WAF records and sampled logs.

Access is not citation. Pages still need clear identities, titles, useful answers, evidence, truthful dates and relevant links. A successful crawl proves neither relevance nor selection, and missing citations cannot be attributed to robots alone.

Configure each purpose separately

Type of accessBusiness decisionTechnical validation
OAI-SearchBotPermit intended SearchBot access and check the effective user-agent group and path rules, rather than assuming every wildcard rule overrides a specific group.Verify official origins, HTTP 200, complete content, canonicals and logs; allow for propagation after policy changes.
GPTBotDecide training preferences through owner approval and applicable rights and compliance policies, independently of search display.Use a separate robots user-agent policy and retain approval and version history.
ChatGPT-UserAssess public pages, authentication, rate limits and security for user-directed visits.Do not treat user-directed visits as automatic search indexing. Test accessible content, semantic controls and interaction errors for relevant browser-agent workflows.
User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

The example expresses 'allow search, disallow potential training'. Replace example domains, sitemaps and paths with real settings. Check the effective robots group and path matching, WAF behavior and approved policy together; rule order alone does not determine precedence.

AI crawler access checklist

  1. Decide the three policies. Have business, technical, legal and site owners approve search, training and user-triggered access boundaries.
  2. Audit effective robots behavior. Check specific and wildcard user-agent groups, path matching, noindex, X-Robots-Tag and sitemaps. Treat robots access and indexing directives as separate controls.
  3. Verify official IP information. Use current OpenAI-published JSON ranges rather than permanently trusting a third-party static list.
  4. Test each delivery layer. Check public robots, home, service, knowledge and evidence URLs for status, TLS, redirects, content and resources.
  5. Inspect CDN and WAF behavior. Check whether CAPTCHA challenges, Bot Fight, geographic blocks, rate limits or User-Agent reputation rules cause intended requests to return HTTP 403 or 429.
  6. Verify server logs. Sample user agents, verified origins, times, URLs, status codes and response sizes; distinguish page retrieval, robots requests and failed retries.
  7. Review public content boundaries. Remove private acceptance records, screenshots, customer identities and unapproved facts. Protect sensitive material with actual access controls.
  8. Retest after changes. The cited OpenAI guidance allows approximately 24 hours for search systems to reflect robots changes. Treat this as propagation guidance, not a crawl or citation deadline; record versions and test times.

Limitations

Allowing OAI-SearchBot only means the applicable robots policy does not opt out. It guarantees neither visits, indexing, rankings, citations nor traffic. User-agent details and official ranges may change. User-triggered ChatGPT-User behavior requires separate testing because robots rules may not apply in the same way.

robots.txt is not a confidentiality mechanism. Use authentication, authorization, network isolation or removal from public hosting for private content. This article is Zhihe Growth's interpretation of documentation accessible on July 22, 2026; production settings require review against actual infrastructure and approved policies.

Official sources