Key Takeaway
On August 31, 2026, Cloudflare announced that Browser Run's `/crawl` endpoint would enforce the `use` restrictions in a site's `robots.txt`. Callers declare reference-level or full use through `contentUse`; a requested use exceeding the site's allowance returns HTTP 400. This provides a documented product-level enforcement example for `use=reference`, but only for this endpoint. Content Signals express preferences and reservations of rights, not universal crawler access controls or citation guarantees.
Facts and background
In its August 31, 2026 announcement, Cloudflare stated that the Browser Run /crawl endpoint reads the target site's robots.txt for the Content Signals use directive. Callers can set contentUse in their requests, with supported values of reference and full in order from more restrictive to more permissive; the default is full. If the site's permitted level is more restrictive than the declared use, Cloudflare rejects the crawl with HTTP 400.
The significant change is product-level enforcement at request time, rather than another declarative robots field. Cloudflare had already defined search,ai-input and ai-train as three purpose signals and was testing the use extension to describe the maximum permitted retention and reuse after access. This update supplies a testable result: if the caller's declared use exceeds the site's allowance, Browser Run's /crawl does not proceed.
Purpose and level of use are separate questions
| Signal or parameter | Question answered | Meaning in the cited documentation | Implication for brands |
|---|---|---|---|
search | Why is the content crawled? | Building a search index with links and short excerpts, excluding AI-generated search summaries | Express a preference about conventional search and link discovery |
ai-input | Why is the content crawled? | Supplying content to a model at query time, including RAG, grounding or generative search answers | Decide separately from search indexing and training permission |
ai-train | Why is the content crawled? | Training or fine-tuning AI models | Search permission does not imply training permission |
use=immediate | How may content be used after access? | Interaction without retention or reuse | The most restrictive use preference |
use=reference | How may content be used after access? | Indexing and excerpts with a link to the source | Public content where attribution is desired and full reuse is restricted |
use=full | How may content be used after access? | Summarization and reproduction | Use only where authorization and business objectives permit |
These two signal groups are not interchangeable. ai-train=no expresses a training restriction, not whether query-time model input or retention is allowed; use=reference limits the level of use without separately defining search, query-time AI input or training purposes. Specify both the purpose of access and permitted subsequent use; copying one example line does not establish a complete policy.
How Cloudflare /crawl enforces these limits
Cloudflare's /crawl documentation also defines the crawlPurposes parameter, allowing callers to declare one or more of search,ai-inputandai-train; the default includes all three. If Content Signals sets any declared purpose to no, the request is rejected with HTTP 400. Separately, contentUse declares the requested use level; a more restrictive site use also causes rejection.
Without explicitly narrowing purpose and use, callers may encounter a conflict between the default of all purposes plus full and the site's policy. For example, a site might allow search, prohibit training and permit only reference; the default request would not comply. Declare the minimum necessary crawlPurposes and contentUse for the actual task. Classify a policy-related 400 as a policy conflict; do not retry blindly or change identity to evade the restriction.
HTTP 400 here is an endpoint enforcement result, not a universal status in every site's bot logs. Other crawlers interpret and enforce Content Signals according to their own products and policies. The absence of Cloudflare /crawl requests does not prove that other AI systems have not crawled or used the content.
Where use=reference helps, and where it does not
For brands seeking search discovery and AI source visibility without permitting full reproduction, use=reference expresses a more specific preference than allowing or blocking everything: indexing and excerpts with source attribution. It may suit public product facts, research, FAQs, evidence records and update logs.
However, reference is not a citation-rate optimization tag. It guarantees neither crawling, indexing, source placement nor traffic. It cannot replace clear titles, truthful update dates, authorship, visible content, canonicals, evidence links and stable URLs.
Nor does it protect quotations, customer records, orders, unreleased products, contracts or internal knowledge. Content Signals cannot replace authentication, authorization, network isolation, WAF controls or data minimization. Deny unauthorized access at the source rather than treating robots.txt as a confidentiality boundary.
Review managed robots.txt defaults individually
The cited Cloudflare documentation states that customers with managed robots.txt enabled receive the combination search=yes, ai-train=no, use=reference. Cloudflare does not automatically set ai-input on their behalf. An omitted signal expresses no preference through that field: it is neither an opt-out nor an affirmative authorization.
Content, legal, growth and security owners should jointly decide ai-input. Public knowledge intended for query-time RAG or grounding needs an explicit authorization decision; conventional search-only content may require a different AI-input policy. Multilingual sites, subdomains and distributor sites should each check the robots.txt actually served by their host, rather than updating only the root domain.
Three common policies and their acceptance evidence
| Content scenario | Candidate policy | Evidence to retain | Not a basis for promising |
|---|---|---|---|
| Public product, knowledge, FAQ and evidence pages | Allow search; decide AI input according to authorization; deny training; consider use=reference | Actual robots response, HTTP 200, complete content, canonical URL, bot logs and source-display retests | Guaranteed AI citations or referral traffic |
| Content authorized for full summaries or reproduction | Use only where contracts and content policies allow use=full, with defined paths and versions | Authorization scope, page inventory, purposes, duration, request records and withdrawal process | That one site-wide permission covers every language, asset or third-party right |
| Customer, transaction, internal or restricted content | Authentication, least privilege, network controls and WAF; do not rely on Content Signals for confidentiality | Denied unauthorized requests, permission logs, data classification and access reviews | That robots declarations stop malicious scraping |
Acceptance testing should cover four layers: whether the raw robots.txt contains the intended signals; whether CDN, cache and host responses agree; whether public target pages consistently return HTTP 200 and complete content; and whether logs and public results separately show the expected crawls, source links and updates. Teams using Browser Run should also test a permitted request with minimum required purposes and a conflicting request expected to return 400. Record policy failures separately from network errors.
Separate policy enforcement from GEO outcomes
The directly testable change is whether Cloudflare /crawl enforces usage boundaries, not whether brand visibility increases. Record signal versions, request parameters, HTTP results, target URLs and times for policy enforcement; bot access, WAF events and extraction for technical access; and source links, cited pages, answer accuracy, mentions and conversions for GEO. These measures are not interchangeable.
If citations fall after setting use=reference, timing alone does not establish causation. Crawl blocks, redesigns, canonical changes, index freshness, platform coverage and question-set variation can all affect results. Keep configuration snapshots and retest fixed questions and pages using logs, citation evidence and site analytics.
Impact on enterprises
First, AI-content governance requires separate decisions about access purpose and subsequent use: search discovery, query-time input, training, retention and reuse cannot be expressed by one blanket Disallow rule. Second, `use=reference` provides a machine-readable preference for discovery with excerpts and attribution, with a documented enforcement example in Cloudflare `/crawl`. This does not establish adoption elsewhere or guarantee a visible source link. Third, Browser Run callers that retain all default `crawlPurposes` and `contentUse=full` may encounter policy-related 400 responses. Diagnose these explicitly rather than masking them with automatic retries. Fourth, govern each host separately. Root domains, language subdomains, distributor sites, help centers and campaign sites may serve robots.txt through different CDNs and templates.
Zhihe Growth's Assessment
The update adds a concrete enforcement example to Content Signals' declarative policy model. Brands should maintain a URL-group policy table covering public access, allowed purposes, maximum use, enforcing controls, logs and public-result checks. Protect nonpublic resources with authentication and authorization first. Then define public-content policies, align robots and CDN/WAF settings, and measure policy enforcement, technical access and GEO outcomes separately. Content authorization is not an SEO trick, and blocking training should not inadvertently block intended search discovery. The August 31 announcement documents `/crawl` enforcement of `use` restrictions and policy-related 400 responses. It does not prove universal standardization or compliance, a particular legal effect, or increased citations and traffic.
Recommended action
- Inventory URL groups for public products, knowledge, FAQs, evidence, campaigns, transactions, customer and internal pages; label their access levels.
- Decide `search`, `ai-input` and `ai-train` separately for each public URL group. Do not treat an omitted field as permission or prohibition.
- Define maximum subsequent use. Evaluate `use=reference` for excerpt and attribution preferences; use `use=full` only with explicit authorization.
- Protect customer records, quotations, orders, contracts and internal knowledge with authentication, authorization, network isolation and WAF controls, not robots.txt alone.
- Check publicly served robots.txt on the root domain, www, language subdomains, help centers and campaign domains; retain dated snapshots.
- Check for conflicts between CDN-managed robots and origin files, then verify consistent responses across regions and protocols after cache refresh.
- For Browser Run, explicitly set the minimum necessary `crawlPurposes` and `contentUse`; do not unknowingly retain all purposes and `full`.
- Classify policy-related HTTP 400 responses separately, recording URLs, parameters and responses. Do not evade restrictions by changing user agents or proxies.
- Monitor bot requests, WAF events, HTTP 200 responses, content completeness, AI source links, citation accuracy and site conversions separately.
- Retest fixed questions and URLs after changes. Retain before-and-after logs and source evidence rather than attributing concurrent fluctuations to one signal.
Limitations
Cloudflare described `use` as an experimental Content Signals extension. The documented enforcement here is Browser Run's `/crawl` endpoint, not every Cloudflare product, AI Search instance, search engine or third-party bot. Other crawlers may ignore unfamiliar robots fields; preference signals are not technical anti-scraping controls. The cited definitions distinguish `reference`, `full` and `immediate`. The endpoint's documented `contentUse` values are `reference` and `full`, while the site-side extension also includes `immediate`. Parameters and defaults may change. Legal effect depends on jurisdiction, contracts, rights and facts; this article is not legal advice. Search permission, AI-input permission and `use=reference` guarantee neither crawl frequency, indexing, rankings, summaries, source placement, recommendations, traffic nor conversions. Protect restricted content with real access controls and substantiate outcome claims with logs, public answers, accessible sources and analytics.