Skip to main content
Platform Developments | AI Bot Governance

Cloudflare's AI Bot Report: Govern Search, Training and Agent Access Separately

On July 1, 2026, Cloudflare reported changes in non-human traffic and training and mixed-use crawlers. The practical question for brands is how to identify access purposes and manage visibility, training permission, security and costs separately.

Key Takeaway

Avoid an undifferentiated block-all-AI rule. Search crawlers, training crawlers, mixed-use bots and user-triggered agents can have different effects. Maintain a joint access matrix covering User-Agent, verifiable IPs, robots rules, CDN/WAF policy and logs, approved by content, legal, security and growth owners.

What's the official change?

A Cloudflare report published on July 1, 2026 said that more than half the traffic on its network came from non-human visitors. Among crawler requests it classified by purpose in June 2026, AI-training crawlers accounted for 52%, mixed-purpose crawlers for more than 36%, and pure search crawlers for a smaller but still important share. These figures describe Cloudflare's network, not an unbiased estimate for the whole internet.

The report argues that combining search, agent use and training under one crawler identity complicates permission decisions. It highlights verifiable identity, declared purpose, discovery signals and licensing. OpenAI's separate descriptions of OAI-SearchBot, GPTBot and ChatGPT-User illustrate the same need to distinguish purposes despite similar operator names.

Date of report 2026-07-01
Key issues Mixed-use bots obscure access purposes
Governance scope robots, WAF, logs, authorizations and costs

Implications for International Brands

International brands may need both public discovery and protection for costly research, paid content or sensitive documents. Blanket AI blocks may reduce discovery; unrestricted access may exceed training, licensing or bandwidth policies. Purpose-specific policies let these goals be evaluated separately.

robots.txt expresses voluntary crawl preferences; it does not replace identity verification, rate limits or application security. User-Agent-only WAF rules can block genuine crawlers or admit spoofed requests. Use the operator's supported verification methods, such as published IP ranges or reverse DNS, alongside behavior and logs, retaining matched rules and response status.

A request with an AI-related User-Agent does not prove a citation, and a citation does not prove a visit or lead. Record logs, answer tests, source URLs, landing sessions and conversions separately to distinguish discovery, training, user visits and unwanted load.

Zhihe Growth's Assessment

Zhihe Growth recommends treating AI bot access as a separate technical acceptance gate. For each bot class, record purpose, official source, access decision, enforcement layer, owner, last verification date and exceptions in the SSOT. Retest after CDN or WAF changes.

For businesses seeking AI-search visibility, search crawlers generally need access. Training access should reflect content value, contracts, applicable requirements and company policy. User-triggered agents require separate boundaries for login, forms, payments and sensitive actions. A static default list cannot govern all three indefinitely.

Implementation Checklist

  1. 1. List official User-Agents, purposes and verification methods for Google, Bing, OpenAI and relevant agents.
  2. 2. Compare robots permissions with actual CDN/WAF rules to detect allowed-in-policy but blocked-at-edge requests.
  3. 3. Sample AI-related server requests and record status, path, IP verification, response-body size and frequency.
  4. 4. Define separate policies for search discovery, training permission, user-triggered access and unknown mixed-use bots.
  5. 5. Set distinct access and rate limits for evidence libraries, downloads, login areas and public knowledge pages.
  6. 6. Retest when platforms change crawler names or IP rules, and update ownership records.

Limitations

Cloudflare's figures reflect its network and classification methods. They do not replace site-specific logs or prove a particular platform trained on, cited or recommended a page.

robots preferences, contracts, copyright and data protection are distinct matters. This technical framework is not legal advice; legal and security owners should assess valuable content and regional requirements.

Official sources