Skip to main content
Industry Insights | Zhihe Growth Research Center

Cloudflare AI Search Discover Crawling: When to Follow Links Instead of Relying on Sitemaps

The source article describes AI Search's discover mode, which can read sitemaps and follow links from a starting point. It compares discover and sitemap modes and covers depth, page limits, caches, robots and path filters.

Key Takeaway

Cloudflare's August 6, 2026 update introduced the `discover` website crawl mode. Starting at the source URL, it can read sitemaps and follow links on crawled pages, helping sites with missing, incomplete or stale sitemaps. It is not a replacement for a well-maintained sitemap: page limits, crawl depth and cache age constrain discovery; it ignores `lastmod` and `changefreq` and does not support specific sitemap selection. The default `sitemap` mode remains preferable when the sitemap is complete and current.

Facts and background

Cloudflare's August 6, 2026 AI Search website-source update introduced the discover parse type. The default sitemap mode reads XML sitemaps declared in robots.txt or configured in the dashboard, without following page links; discover starts a crawl at the source URL and, by default, finds candidate URLs through both sitemaps and page links. Retrieved pages are stored, converted to Markdown, chunked and indexed in AI Search.

Discovery for an enterprise knowledge index, not public search indexing

Cloudflare AI Search provides retrieval infrastructure for applications, customer support, distributor portals and agents. discover determines how that index discovers pages on a website in the same Cloudflare account. It does not replace Google, Bing or ChatGPT crawling, ranking or citation systems, or automatically increase public AI-search visibility.

Key differences between the modes

Decisionsitemap (default)discover
Page discoveryReads configured sitemaps or those declared in robots.txt; does not follow linksReads sitemaps and follows page links by default
Best suited toA complete, current sitemapA missing, incomplete or stale sitemap
Update signalReads lastmod; falls back to changefreqIgnores lastmod and changefreq; uses max_age to determine when to refetch
Coverage limitsIndexes pages listed in the sitemapConstrained by limit,depth and the instance object limit
Indexing orderUses sitemap priority to determine indexing orderDoes not use sitemap priority to control order
Specific sitemapsUp to five can be configuredNot supportedparse_options.specific_sitemaps
Switching modes—Changing an existing instance's parse type triggers a full resync

Discover fills gaps but does not repair information architecture

discoverThe documented defaults are 100,000 pages and a link depth of 5, adjustable within the supported ranges. Discover can find linked content omitted from a sitemap, but deep or isolated pages, pages reachable only through site search, incorrect canonicals and poorly linked market versions may still be missed. If a crawl reaches its page limit or does not finish, Cloudflare retains previously indexed pages that the crawl did not reach. A successful job therefore does not establish complete coverage.

Freshness shifts from lastmod to cache age

sitemap mode recrawls a page when lastmod is later than the previous sync; discover instead uses max_age to control how long cached content can be reused. The documented default is 86,400 seconds, with a range of 0 to 604,800 seconds. Setting 0 fetches from the origin each time, potentially increasing origin load and crawl costs. Frequently changing prices, stock, delivery coverage and policies require coordinated cache settings, sync schedules and publishing triggers, not just a visible update date.

Crawler access and content scope still matter

On domains in the same account, the crawler uses the Cloudflare-AI-Search user agent. If external links or subdomains are enabled and the crawl reaches a domain outside the account, it uses Cloudflare-AI-Search-External. Pages disallowed by robots.txt are recorded as blocked_by_robots_txt. Cloudflare also documents that the crawler declares the search and ai-input purposes; if Content Signals sets either to no, the crawl is rejected. Even with access allowed, path filters, content selectors, authentication and WAF rules can leave the index with error pages, navigation noise or incomplete content.

Impact on enterprises

Multilingual, multi-market and multi-CMS sites often have sitemaps but inconsistent discovery of product, policy, evidence or regional pages. Discover can fill gaps while also exposing architectural weaknesses: deep links cause omissions, unrestricted paths admit login, parameter, filter or obsolete pages, and external or subdomain crawling can greatly expand scope. Choosing a crawl mode is a knowledge-governance decision, not merely a dashboard setting.

Zhihe Growth's Assessment

Repair sitemaps and internal links first, then decide whether Discover is needed. Keep the default mode when the site has a complete, auditable sitemap with truthful `lastmod` values; use link discovery to audit coverage, not replace the authoritative URL list. Consider `discover` for legacy CMS limitations, temporary migration gaps or link-organized knowledge sources. Define allowed paths, page limits, depth, cache age and a full-rebuild window before launch.

Recommended action

  1. Compare sitemap URLs with crawlable internal links, published CMS records and canonical URLs to identify coverage gaps.
  2. Keep the default `sitemap` mode when coverage is complete and `lastmod` is reliable; do not switch solely to crawl more pages.
  3. For `discover`, define the source URL, discovery source (`sitemaps`, `links` or `all`), `limit`, `depth` and `max_age`. Record that switching modes triggers a full resync.
  4. Exclude login, administration, site-search, parameterized filter, preview, draft, archive-copy and unauthorized paths. Check case sensitivity and trailing slashes; exclusion rules take precedence.
  5. Check robots.txt and Content Signals for the `Cloudflare-AI-Search` crawler's `search` and `ai-input` purposes. Verify `Cloudflare-AI-Search-External` separately when crawling domains outside the account.
  6. Aim to place product, policy, evidence and guide pages within two visible internal-link hops. Do not rely solely on site-search forms or JavaScript events to expose isolated pages.
  7. Sample final URLs, canonicals, HTTP 200 responses, complete content, market and language versions, update dates and source links; do not inspect only the indexed count.
  8. Record CMS releases, AI Search syncs, failures, sampled answers and rollback steps in the update log. Do not accept a release as fully covered when page or depth limits are reached or a full sync is unfinished.

Limitations

The original article described AI Search as Beta; features, limits and interfaces may change. Website sources require a source domain onboarded to the same Cloudflare account. Discover is limited to 100,000 pages, with actual coverage also constrained by instance limits, depth, filters, robots.txt, Content Signals, WAF, rendering and cache settings. Indexed status does not establish complete retrieval or correct answers. AI Search configuration does not guarantee public Google, Bing or ChatGPT indexing, rankings, citations, traffic or conversions.

Source Verification