Key Takeaway
Cloudflare's August 6, 2026 update introduced the `discover` website crawl mode. Starting at the source URL, it can read sitemaps and follow links on crawled pages, helping sites with missing, incomplete or stale sitemaps. It is not a replacement for a well-maintained sitemap: page limits, crawl depth and cache age constrain discovery; it ignores `lastmod` and `changefreq` and does not support specific sitemap selection. The default `sitemap` mode remains preferable when the sitemap is complete and current.
Facts and background
Cloudflare's August 6, 2026 AI Search website-source update introduced the discover parse type. The default sitemap mode reads XML sitemaps declared in robots.txt or configured in the dashboard, without following page links; discover starts a crawl at the source URL and, by default, finds candidate URLs through both sitemaps and page links. Retrieved pages are stored, converted to Markdown, chunked and indexed in AI Search.
Discovery for an enterprise knowledge index, not public search indexing
Cloudflare AI Search provides retrieval infrastructure for applications, customer support, distributor portals and agents. discover determines how that index discovers pages on a website in the same Cloudflare account. It does not replace Google, Bing or ChatGPT crawling, ranking or citation systems, or automatically increase public AI-search visibility.
Key differences between the modes
| Decision | sitemap (default) | discover |
|---|---|---|
| Page discovery | Reads configured sitemaps or those declared in robots.txt; does not follow links | Reads sitemaps and follows page links by default |
| Best suited to | A complete, current sitemap | A missing, incomplete or stale sitemap |
| Update signal | Reads lastmod; falls back to changefreq | Ignores lastmod and changefreq; uses max_age to determine when to refetch |
| Coverage limits | Indexes pages listed in the sitemap | Constrained by limit,depth and the instance object limit |
| Indexing order | Uses sitemap priority to determine indexing order | Does not use sitemap priority to control order |
| Specific sitemaps | Up to five can be configured | Not supportedparse_options.specific_sitemaps |
| Switching modes | — | Changing an existing instance's parse type triggers a full resync |
Discover fills gaps but does not repair information architecture
discoverThe documented defaults are 100,000 pages and a link depth of 5, adjustable within the supported ranges. Discover can find linked content omitted from a sitemap, but deep or isolated pages, pages reachable only through site search, incorrect canonicals and poorly linked market versions may still be missed. If a crawl reaches its page limit or does not finish, Cloudflare retains previously indexed pages that the crawl did not reach. A successful job therefore does not establish complete coverage.
Freshness shifts from lastmod to cache age
sitemap mode recrawls a page when lastmod is later than the previous sync; discover instead uses max_age to control how long cached content can be reused. The documented default is 86,400 seconds, with a range of 0 to 604,800 seconds. Setting 0 fetches from the origin each time, potentially increasing origin load and crawl costs. Frequently changing prices, stock, delivery coverage and policies require coordinated cache settings, sync schedules and publishing triggers, not just a visible update date.
Crawler access and content scope still matter
On domains in the same account, the crawler uses the Cloudflare-AI-Search user agent. If external links or subdomains are enabled and the crawl reaches a domain outside the account, it uses Cloudflare-AI-Search-External. Pages disallowed by robots.txt are recorded as blocked_by_robots_txt. Cloudflare also documents that the crawler declares the search and ai-input purposes; if Content Signals sets either to no, the crawl is rejected. Even with access allowed, path filters, content selectors, authentication and WAF rules can leave the index with error pages, navigation noise or incomplete content.
Impact on enterprises
Multilingual, multi-market and multi-CMS sites often have sitemaps but inconsistent discovery of product, policy, evidence or regional pages. Discover can fill gaps while also exposing architectural weaknesses: deep links cause omissions, unrestricted paths admit login, parameter, filter or obsolete pages, and external or subdomain crawling can greatly expand scope. Choosing a crawl mode is a knowledge-governance decision, not merely a dashboard setting.
Zhihe Growth's Assessment
Repair sitemaps and internal links first, then decide whether Discover is needed. Keep the default mode when the site has a complete, auditable sitemap with truthful `lastmod` values; use link discovery to audit coverage, not replace the authoritative URL list. Consider `discover` for legacy CMS limitations, temporary migration gaps or link-organized knowledge sources. Define allowed paths, page limits, depth, cache age and a full-rebuild window before launch.
Recommended action
- Compare sitemap URLs with crawlable internal links, published CMS records and canonical URLs to identify coverage gaps.
- Keep the default `sitemap` mode when coverage is complete and `lastmod` is reliable; do not switch solely to crawl more pages.
- For `discover`, define the source URL, discovery source (`sitemaps`, `links` or `all`), `limit`, `depth` and `max_age`. Record that switching modes triggers a full resync.
- Exclude login, administration, site-search, parameterized filter, preview, draft, archive-copy and unauthorized paths. Check case sensitivity and trailing slashes; exclusion rules take precedence.
- Check robots.txt and Content Signals for the `Cloudflare-AI-Search` crawler's `search` and `ai-input` purposes. Verify `Cloudflare-AI-Search-External` separately when crawling domains outside the account.
- Aim to place product, policy, evidence and guide pages within two visible internal-link hops. Do not rely solely on site-search forms or JavaScript events to expose isolated pages.
- Sample final URLs, canonicals, HTTP 200 responses, complete content, market and language versions, update dates and source links; do not inspect only the indexed count.
- Record CMS releases, AI Search syncs, failures, sampled answers and rollback steps in the update log. Do not accept a release as fully covered when page or depth limits are reached or a full sync is unfinished.
Limitations
The original article described AI Search as Beta; features, limits and interfaces may change. Website sources require a source domain onboarded to the same Cloudflare account. Discover is limited to 100,000 pages, with actual coverage also constrained by instance limits, depth, filters, robots.txt, Content Signals, WAF, rendering and cache settings. Indexed status does not establish complete retrieval or correct answers. AI Search configuration does not guarantee public Google, Bing or ChatGPT indexing, rankings, citations, traffic or conversions.