Crawler permissions are a publishing-governance decision, not simply a matter of filling robots.txt with a series of Allow: /. Businesses must answer four separate questions: should public pages enter search indexes, may AI search discover them, may crawlers collect them for model training, and can they be accessed when a user supplies a link directly? These questions have different business implications and cannot be reduced to one master switch.

Classify Access by Purpose, Not Just the AI Label

Visitors Typical uses Policy to Confirm Separately
Googlebot Google Search crawling and indexing Whether public pages should appear in search
OAI-SearchBot Discovery for OpenAI search Whether the site should appear in relevant search results
GPTBot Crawling for OpenAI model training Whether content may be used for training
ChatGPT-User User-initiated page access Whether public pages can be read in real time

OpenAI official crawler documentation distinguishes these purposes and provides independent controls. Allowing GPTBot is not a prerequisite for inclusion in ChatGPT search, and a user-initiated visit is not evidence of periodic indexing. Check other platforms' own documentation rather than copying OpenAI's crawler tokens.

The Separate Roles of Robots Rules, Indexing, and Security Controls

robots.txt provides access instructions for compliant crawlers; it is not an authentication or confidentiality tool. Search systems may still discover a blocked URL through external links. Google's robots guidance also explains that keeping a page out of results requires controls suited to the goal, such as noindex or login protection. Sensitive information must not be protected by robots rules alone. noindex itself requires the crawler to read the page to see the directive, so blocking crawling in robots while placing noindex on the page is often a contradictory combination.

Actual requests also pass through DNS, TLS, CDN/WAF controls, origin routing, and application permissions. Even when robots rules allow access, a WAF 403 response, challenge page, or regional block can make the body unreadable. Conversely, anonymous browser access does not prove that a particular crawler receives the same response. Public-content policies must be checked at the protocol, edge, and origin layers. Administrative interfaces, customer data, and unpublished content require real access controls.

An Actionable Access Matrix

During a GEO technical assessment, Zhihe Growth can classify URLs as public content, public resources, private customer content, or administrative endpoints. For each crawler, record expected and observed robots decisions, HTTP status, response type, body availability, challenges, and the last successful crawl. Have the client approve training and search permissions before changes, then verify server logs and platform tools afterward. Keep simulated or proxy requests distinct from genuine crawler results.

For example, allowing search access to service pages and prohibiting anonymous access to customer exports are separate rules. Start with CDN/WAF and HTTP Status Diagnostics, then check Log Verification Methods, and finally use Citation Rate Measurement to observe outcomes. The full sequence is access permission -> crawling -> indexing/retrieval -> answer -> citation. Each step requires its own evidence.

Acceptance Checks and When to Revoke Access

Acceptance requires more than inspecting robots.txt. Sample key URLs, verify actual status codes and body content, check CAPTCHAs and redirects, compare edge and origin logs, verify crawler identities, and then inspect indexing status in the platform console. If customer information is unexpectedly exposed, revoke public access and address the disclosure first. Do not keep it accessible for GEO. Site owners approve crawling permissions; Zhihe Growth provides assessments, implementation advice, and retesting, without treating permission as a guarantee of indexing or AI citation.

How to Write a Change Request

Suppose a client wants public knowledge articles to appear in ChatGPT search but does not permit training crawls. Record the target URL patterns, disclosure level, and approver in the change request, then configure OAI-SearchBot and GPTBot separately. Verify that robots.txt is accessible on the correct domain, protocol, and port. If the site has both www and an apex domain, check the policy on each final destination domain rather than relying on a Saved message in the admin console. Sample service pages, articles, PDFs, and customer directories, comparing expected and actual responses. A WAF allowlist must not rely solely on a self-reported User-Agent, which any script can imitate. Use the platform's official verification method together with edge logs.

Record results in four columns: page type, permission policy, observed response, and remaining risks. If a public article returns 200 but carries noindex, crawlability and indexability remain separate. If private customer data returns 200 to an unauthorized request, address access controls immediately. Only after testing should you update URL Governance and the sitemap. This sequence avoids the irreversible risk of exposing the entire site before adding security rules.

Return to Knowledge Center · Explore GEO Services