Key Takeaway
Cloudflare's July 1, 2026 announcement scheduled new-domain defaults for September 15: Search allowed, with Agent and Training blocked on pages Cloudflare identifies as displaying ads. Mixed Search-and-Training crawlers are also affected by training-block rules. This is not a blanket ban on all AI traffic. Before onboarding or migration, check all three policies, ad-page classification and existing WAF rules, then validate actual status codes, crawl paths and referrals.
Facts and background
This article was prepared on September 13, 2026, ahead of an announced implementation date. Cloudflare had announced the change on July 1, with updated defaults for newly onboarded domains from September 15. It was not a feature first announced on September 13, and the announcement did not establish that all existing domains would automatically receive identical policies. The source's assessment of the preceding 72 hours was an editorial selection judgment, not an exhaustive platform changelog.
The defaults distinguish behavior rather than treating all AI crawlers alike. Search collects or indexes content for later answers; Agent covers real-time user-directed activity such as chat fetches and browser tasks; Training covers model training or fine-tuning. A bot can have multiple behaviors, so neither its operator name nor its User-Agent alone determines access.
What the September 15 Change Covered
| Component | Announced default or rule change | What brands should confirm | What cannot be inferred |
|---|---|---|---|
| Newly onboarded Cloudflare domains | Search allowed; Agent and Training blocked by default on pages displaying ads | Onboarding date, current settings for all three behaviors and detected ad-supported pages | That all AI access to the entire site is blocked |
| Mixed-use crawlers | Crawlers combining Search and Training are affected by training-block settings, including the legacy Block AI bots configuration described in the source | All behaviors currently assigned to the crawler and the rule that actually determines access | That any Search classification guarantees access |
| Legacy Block AI bots setting | The source documentation marked it for deprecation on September 15; earlier guidance excluded mixed-use crawlers | Whether settings have migrated to separate Search, Agent and Training policies | That the old switch and new policies remain equivalent |
| Existing domains | The announcement addresses new domains and allowed customers to opt out of the new defaults before the effective date | Actual console settings, custom rules and change records | That domain creation time alone proves live behavior |
Block on pages with ads uses Cloudflare's automatic detection, not a fixed list supplied by the site owner. Test ad templates, regional differences, A/B variants and dynamic loading. Ad-free product pages, ad-supported articles and checkout may behave differently; a homepage-only test is insufficient.
Mixed-use classification is an important boundary. The cited guidance says a bot may have multiple behaviors, with Search-and-Training crawlers affected by training blocks from September 15. Checking only Search=Allow can miss a Training classification that produces HTTP 403 or another block on ad-supported pages. Conversely, allowing an operator's entire traffic may exceed a search-only permission policy.
robots.txt, AI Bot Policies and WAF Are Separate Layers
| Control Layer | Primary purpose | Technically enforced? | Validation focus |
|---|---|---|---|
| robots.txt and Content Signals | Express crawl and content-use preferences, such as search, ai-input, ai-train or use | robots.txt compliance is voluntary; Content Signals are not a security boundary | Published rules, their intended crawlers and evidence of reading and compliance |
| AI Bot Behavior Policy | Choose Allow, site-wide Block or ad-page-only Block for Search, Agent and Training | Cloudflare executes through its edge control. | Actual hit category, page, status code and rule action |
| Individual crawler controls | Allow or block a specific crawler | Enforced through AI Crawl Control | Consistency of operator, verified crawler identity, classification and exceptions |
| WAF Custom Rules | Apply granular conditions by path, hostname, detection field or other criteria | Technically enforced before later bot-processing and Pay Per Crawl stages described in the cited guidance | Rule order, skip rules, false blocks and overlaps |
Cloudflare describes robots.txt as a preference mechanism, not a technical barrier to noncompliant crawlers. Use enforced controls where required. If managed robots.txt, legacy bot blocks, AI Crawl Control and custom WAF rules coexist, trace a real request through them to identify the final action. A remaining HTTP 403 after changing one setting does not by itself establish a platform fault.
The cited Cloudflare documentation says AI Crawl Control blocking uses WAF custom rules and occurs before subsequent Bot Solutions and Pay Per Crawl processing. Audit existing WAF rules before changing AI bot policies: an earlier generic bot, regional or path rule can still block a Search crawler even when the policy interface shows Allow.
Do Not Validate Identity from User-Agent Alone
Detection varies by plan. The cited introduction described Free-plan detection primarily through known self-declared User-Agents, with richer detection IDs available in Enterprise Bot Management. User-Agents can be spoofed; a simulated request proves only the response to that string, not the operator's identity.
Retain four evidence types: configuration snapshots; behavior classifications from BotBase or the current crawler list; edge logs with detection fields, matched rules and status codes; and actual platform crawl or referral records. Test representative Search, Agent and Training crawlers across ad-free core pages, ad-supported articles, product pages, robots.txt, sitemaps and a restricted path.
| Test Results | Decision | Next step |
|---|---|---|
| Search receives 2xx on core public pages; Agent and Training are blocked only where intended; WAF and robots policies agree | Pass | Record configuration versions to continuously monitor changes in new classifications and page templates |
| Policies match expectations, but ad detection or regional templates produce inconsistent results | Conditional pass | Limit rollout and retest templates, regions and dynamic ad loading |
| Legacy WAF rules, training policy or mixed-use classification unexpectedly block Search | Fail | Pause migration or publication, identify the matched rule and retest |
| Only simulated User-Agent results are available, without detection IDs, logs or actual platform evidence | Cannot establish attribution | Do not claim real AI-search crawler access is established; obtain identity and live-request evidence |
Measure Policy and Business Outcomes Separately
The cited AI Crawl Control guidance describes request analysis by date, crawler, operator, hostname and path, including allowed requests, failures, status codes, transfer volume and popular paths. Some plans provide AI referral information. Establish who requested what and which rule applied before evaluating referrals, indexing, citations or conversions. A 2xx crawl response is access evidence, not proof of business results.
Manage multi-domain and multilingual sites at zone and hostname level. Newly onboarded campaign domains, regional subdomains, migrated domains and landing pages may have different defaults and WAF histories. Start each onboarding with a baseline, configuration comparison and real-path sampling, not account-wide assumptions.
Impact on enterprises
The announced change moves access policy beyond a single AI-bot switch toward behavior, page monetization and detected identity. Allowing Search is a useful default direction for brands seeking discovery, but mixed-use classifications, ad detection and existing WAF rules can alter actual responses. The main risk is not knowing which layer enforces a decision. Marketing may need search referrals, legal teams may restrict training and product teams may need user-directed access to public product facts. A blanket rule can obscure these distinct requirements and reduce both access and auditability.
Zhihe Growth's Assessment
Treat the September 15 defaults as an access-policy migration. Decide permitted content uses, map Search, Agent and Training to page types, then reconcile robots.txt, managed policies, WAF and crawler exceptions. Validate actual ad-page classifications instead of reading settings alone. Keep search access separate from training permission, and do not adopt platform defaults as final company policy without review. Assess agent access separately for public quotes, stock, store and support information. Protect accounts, payments, private prices and customer data through authentication and application security. This article describes the product rules cited as of September 13, 2026. It does not establish correct classification of any particular page or guarantee crawling, indexing, citations, referrals or conversions.
Recommended action
- Inventory zones, subdomains and campaign domains onboarded around September 15, recording dates, owners and current AI bot policies.
- Export or capture Search, Agent and Training settings separately; do not substitute the legacy Block AI bots switch for policy reconciliation.
- Check every behavior assigned to priority crawlers in BotBase or the current list, flagging combined Search-and-Training crawlers.
- Test ad-free core pages, ad-supported articles, product pages, robots.txt, sitemaps and restricted paths across principal regions and templates.
- Audit custom WAF rules, Bot Fight Mode, AI Crawl Control and crawler exceptions for ordering and conflicting actions.
- Treat robots.txt and Content Signals as preferences and AI bot policies and WAF as enforcement. Save configuration and response evidence separately.
- Do not identify real crawlers from simulated User-Agents alone. Use available detection IDs and security logs to support attribution.
- After launch or migration, inspect 2xx, 3xx, 4xx and 5xx responses by crawler, operator, hostname and path to locate unintended blocks.
- Report access, crawling, indexing, citations, referrals and conversions separately. Allowed-request counts are not AI-search performance results.
- Review changes to policies, bot classifications and ad detection. Resample after template, ad-system, domain-onboarding or classification changes.
Limitations
This article uses Cloudflare documentation cited as available on September 13, 2026. July 1 was the announcement date and September 15 the announced effective date, not a feature launched in the preceding 72 hours. The stated defaults concern newly onboarded domains; establish existing-domain behavior from configuration, WAF rules and live requests. Search, Agent and Training are Cloudflare behavior categories, not equivalents of every platform's crawler names, robots policies or data-use commitments. Categories can overlap and change. User-Agent detection has spoofing and attribution limits, and detection IDs do not replace rule and response checks. Cloudflare detects ad-supported pages automatically. This article did not inspect a customer's console and guarantees no page classification. robots.txt and Content Signals do not provide authentication, privacy or data-loss prevention. Protect sensitive paths through access controls and application security. Allowing Search establishes only part of edge-access policy, not crawling, indexing, presentation, citations or business outcomes.