Skip to main content
Industry Insights | Zhihe Growth Research Center

Cloudflare's September 15 AI Bot Defaults: How Can Brands Avoid Blocking AI Search?

Cloudflare's July 1, 2026 announcement set September 15 as the start date for new-domain defaults: Search allowed, with Agent and Training blocked on ad-supported pages. This guide covers scope, mixed-use crawlers, robots.txt, WAF and validation.

Key Takeaway

Cloudflare's July 1, 2026 announcement scheduled new-domain defaults for September 15: Search allowed, with Agent and Training blocked on pages Cloudflare identifies as displaying ads. Mixed Search-and-Training crawlers are also affected by training-block rules. This is not a blanket ban on all AI traffic. Before onboarding or migration, check all three policies, ad-page classification and existing WAF rules, then validate actual status codes, crawl paths and referrals.

Facts and background

This article was prepared on September 13, 2026, ahead of an announced implementation date. Cloudflare had announced the change on July 1, with updated defaults for newly onboarded domains from September 15. It was not a feature first announced on September 13, and the announcement did not establish that all existing domains would automatically receive identical policies. The source's assessment of the preceding 72 hours was an editorial selection judgment, not an exhaustive platform changelog.

The defaults distinguish behavior rather than treating all AI crawlers alike. Search collects or indexes content for later answers; Agent covers real-time user-directed activity such as chat fetches and browser tasks; Training covers model training or fine-tuning. A bot can have multiple behaviors, so neither its operator name nor its User-Agent alone determines access.

What the September 15 Change Covered

ComponentAnnounced default or rule changeWhat brands should confirmWhat cannot be inferred
Newly onboarded Cloudflare domainsSearch allowed; Agent and Training blocked by default on pages displaying adsOnboarding date, current settings for all three behaviors and detected ad-supported pagesThat all AI access to the entire site is blocked
Mixed-use crawlersCrawlers combining Search and Training are affected by training-block settings, including the legacy Block AI bots configuration described in the sourceAll behaviors currently assigned to the crawler and the rule that actually determines accessThat any Search classification guarantees access
Legacy Block AI bots settingThe source documentation marked it for deprecation on September 15; earlier guidance excluded mixed-use crawlersWhether settings have migrated to separate Search, Agent and Training policiesThat the old switch and new policies remain equivalent
Existing domainsThe announcement addresses new domains and allowed customers to opt out of the new defaults before the effective dateActual console settings, custom rules and change recordsThat domain creation time alone proves live behavior

Block on pages with ads uses Cloudflare's automatic detection, not a fixed list supplied by the site owner. Test ad templates, regional differences, A/B variants and dynamic loading. Ad-free product pages, ad-supported articles and checkout may behave differently; a homepage-only test is insufficient.

Mixed-use classification is an important boundary. The cited guidance says a bot may have multiple behaviors, with Search-and-Training crawlers affected by training blocks from September 15. Checking only Search=Allow can miss a Training classification that produces HTTP 403 or another block on ad-supported pages. Conversely, allowing an operator's entire traffic may exceed a search-only permission policy.

robots.txt, AI Bot Policies and WAF Are Separate Layers

Control LayerPrimary purposeTechnically enforced?Validation focus
robots.txt and Content SignalsExpress crawl and content-use preferences, such as search, ai-input, ai-train or userobots.txt compliance is voluntary; Content Signals are not a security boundaryPublished rules, their intended crawlers and evidence of reading and compliance
AI Bot Behavior PolicyChoose Allow, site-wide Block or ad-page-only Block for Search, Agent and TrainingCloudflare executes through its edge control.Actual hit category, page, status code and rule action
Individual crawler controlsAllow or block a specific crawlerEnforced through AI Crawl ControlConsistency of operator, verified crawler identity, classification and exceptions
WAF Custom RulesApply granular conditions by path, hostname, detection field or other criteriaTechnically enforced before later bot-processing and Pay Per Crawl stages described in the cited guidanceRule order, skip rules, false blocks and overlaps

Cloudflare describes robots.txt as a preference mechanism, not a technical barrier to noncompliant crawlers. Use enforced controls where required. If managed robots.txt, legacy bot blocks, AI Crawl Control and custom WAF rules coexist, trace a real request through them to identify the final action. A remaining HTTP 403 after changing one setting does not by itself establish a platform fault.

The cited Cloudflare documentation says AI Crawl Control blocking uses WAF custom rules and occurs before subsequent Bot Solutions and Pay Per Crawl processing. Audit existing WAF rules before changing AI bot policies: an earlier generic bot, regional or path rule can still block a Search crawler even when the policy interface shows Allow.

Do Not Validate Identity from User-Agent Alone

Detection varies by plan. The cited introduction described Free-plan detection primarily through known self-declared User-Agents, with richer detection IDs available in Enterprise Bot Management. User-Agents can be spoofed; a simulated request proves only the response to that string, not the operator's identity.

Retain four evidence types: configuration snapshots; behavior classifications from BotBase or the current crawler list; edge logs with detection fields, matched rules and status codes; and actual platform crawl or referral records. Test representative Search, Agent and Training crawlers across ad-free core pages, ad-supported articles, product pages, robots.txt, sitemaps and a restricted path.

Test ResultsDecisionNext step
Search receives 2xx on core public pages; Agent and Training are blocked only where intended; WAF and robots policies agreePassRecord configuration versions to continuously monitor changes in new classifications and page templates
Policies match expectations, but ad detection or regional templates produce inconsistent resultsConditional passLimit rollout and retest templates, regions and dynamic ad loading
Legacy WAF rules, training policy or mixed-use classification unexpectedly block SearchFailPause migration or publication, identify the matched rule and retest
Only simulated User-Agent results are available, without detection IDs, logs or actual platform evidenceCannot establish attributionDo not claim real AI-search crawler access is established; obtain identity and live-request evidence

Measure Policy and Business Outcomes Separately

The cited AI Crawl Control guidance describes request analysis by date, crawler, operator, hostname and path, including allowed requests, failures, status codes, transfer volume and popular paths. Some plans provide AI referral information. Establish who requested what and which rule applied before evaluating referrals, indexing, citations or conversions. A 2xx crawl response is access evidence, not proof of business results.

Manage multi-domain and multilingual sites at zone and hostname level. Newly onboarded campaign domains, regional subdomains, migrated domains and landing pages may have different defaults and WAF histories. Start each onboarding with a baseline, configuration comparison and real-path sampling, not account-wide assumptions.

Impact on enterprises

The announced change moves access policy beyond a single AI-bot switch toward behavior, page monetization and detected identity. Allowing Search is a useful default direction for brands seeking discovery, but mixed-use classifications, ad detection and existing WAF rules can alter actual responses. The main risk is not knowing which layer enforces a decision. Marketing may need search referrals, legal teams may restrict training and product teams may need user-directed access to public product facts. A blanket rule can obscure these distinct requirements and reduce both access and auditability.

Zhihe Growth's Assessment

Treat the September 15 defaults as an access-policy migration. Decide permitted content uses, map Search, Agent and Training to page types, then reconcile robots.txt, managed policies, WAF and crawler exceptions. Validate actual ad-page classifications instead of reading settings alone. Keep search access separate from training permission, and do not adopt platform defaults as final company policy without review. Assess agent access separately for public quotes, stock, store and support information. Protect accounts, payments, private prices and customer data through authentication and application security. This article describes the product rules cited as of September 13, 2026. It does not establish correct classification of any particular page or guarantee crawling, indexing, citations, referrals or conversions.

Recommended action

  1. Inventory zones, subdomains and campaign domains onboarded around September 15, recording dates, owners and current AI bot policies.
  2. Export or capture Search, Agent and Training settings separately; do not substitute the legacy Block AI bots switch for policy reconciliation.
  3. Check every behavior assigned to priority crawlers in BotBase or the current list, flagging combined Search-and-Training crawlers.
  4. Test ad-free core pages, ad-supported articles, product pages, robots.txt, sitemaps and restricted paths across principal regions and templates.
  5. Audit custom WAF rules, Bot Fight Mode, AI Crawl Control and crawler exceptions for ordering and conflicting actions.
  6. Treat robots.txt and Content Signals as preferences and AI bot policies and WAF as enforcement. Save configuration and response evidence separately.
  7. Do not identify real crawlers from simulated User-Agents alone. Use available detection IDs and security logs to support attribution.
  8. After launch or migration, inspect 2xx, 3xx, 4xx and 5xx responses by crawler, operator, hostname and path to locate unintended blocks.
  9. Report access, crawling, indexing, citations, referrals and conversions separately. Allowed-request counts are not AI-search performance results.
  10. Review changes to policies, bot classifications and ad detection. Resample after template, ad-system, domain-onboarding or classification changes.

Limitations

This article uses Cloudflare documentation cited as available on September 13, 2026. July 1 was the announcement date and September 15 the announced effective date, not a feature launched in the preceding 72 hours. The stated defaults concern newly onboarded domains; establish existing-domain behavior from configuration, WAF rules and live requests. Search, Agent and Training are Cloudflare behavior categories, not equivalents of every platform's crawler names, robots policies or data-use commitments. Categories can overlap and change. User-Agent detection has spoofing and attribution limits, and detection IDs do not replace rule and response checks. Cloudflare detects ad-supported pages automatically. This article did not inspect a customer's console and guarantees no page classification. robots.txt and Content Signals do not provide authentication, privacy or data-loss prevention. Protect sensitive paths through access controls and application security. Allowing Search establishes only part of edge-access policy, not crawling, indexing, presentation, citations or business outcomes.

Source Verification