Key Takeaway
Cloudflare now lets Browser Run `/crawl` jobs push start, per-page status update, and finish events to Queues, instead of relying only on periodic status polling. International brands that use web content for their own AI search, support knowledge bases, or RAG systems should use event subscriptions for scheduling, failure alerts, and auditing, then actively retrieve complete results by job ID after completion. The events contain no page content and do not prove that Google, Bing, ChatGPT, or other third-party platforms have crawled, indexed, or cited the brand's content.
Facts and Background
On September 25, 2026, Cloudflare announced that Browser Run /crawl jobs can publish lifecycle events to Cloudflare Queues. The three documented event types are crawl.started,crawl.updated and crawl.finished. They indicate that a job has started, a discovered URL has changed status, or the overall job has finished. Cloudflare positions subscriptions as a push-based alternative to status polling, for tracking progress or triggering downstream processing.
This changes a company's own web collection and knowledge-update workflow, not public search engines' indexing protocols. Events show collection progress, which URLs succeeded or failed, and when a job ended. They do not show whether a third-party AI platform has indexed content, how it understands it, or when it will cite it. Assess internal knowledge-base ingestion, public search access, and final citation performance separately.
What the Three Event Types Establish
| Event | Key fields | Suitable actions to trigger | What the event alone does not prove |
|---|---|---|---|
crawl.started | jobId,createdAt,crawlConfig | Register the job version, lock the configuration, and start the timeout clock | That any page has been crawled successfully |
crawl.updated | jobId,URL,crawlStatus,httpStatus | Update per-page records, flag 403/404/5xx responses, and trigger targeted alerts | That page content is in the vector store or can support correct answers |
crawl.finished | jobStatus, start and end times, total, completed, errored, skipped, crawlConfig | Reconcile totals, fetch complete results, and decide whether to proceed to parsing and indexing | That all target URLs are covered or the knowledge base has been published |
Browser Run events are account-level: one subscription receives events for every crawl job in the account without a source-specific selector. A consumer cannot assume its queue serves only one domain, market, or knowledge base. At minimum, join the job ID to a registry containing the target domain, environment, language, market, content purpose, configuration version, and responsible owner. This keeps test, production, and different brands' data streams separate.
Event metadata also includes accountId, eventSubscriptionId, eventSchemaVersion, and eventTimestamp. Cloudflare Queues provides at-least-once delivery by default, so the same message may occasionally arrive more than once. Use idempotency keys and make duplicate events safe to replay. Two update deliveries do not mean a page was crawled twice and must not cause duplicate indexing or publication.
Events Are Not Content: Retrieve Complete Results After Completion
Cloudflare explicitly states that lifecycle events carry status information, not crawled page content. After receiving crawl.finished, use the job ID to call the results endpoint, retrieve every page of records, and handle URL statuses such as completed,errored,disallowed,skipped or cancelled separately. Responses exceeding 10 MB include a cursor. Do not mark content as fully ingested until all result pages have been retrieved.
Current documentation allows crawl jobs to run for up to 7 days and retains results for 14 days after completion. Saving events without retrieving and archiving the results within that window risks losing page content and per-page diagnostic evidence. A safer workflow moves a finished job only to 'awaiting retrieval.' Mark the knowledge base as updated only after complete retrieval, validation, parsing, deduplication, permission filtering, and indexing succeed.
Use a Five-Stage State Machine, Not One Completion Flag
| Stage | Entry condition | Required evidence | Failure action |
|---|---|---|---|
| Job registered | Create the job and obtain its job ID | Target URLs, configuration hash, market, language, and initiation time | Do not enqueue downstream work |
| Crawl running | Receive started or query a running status | crawlConfig, event time, and timeout deadline | Alert on timeout without assuming a cause |
| Per-page validation | Receive updated | URL, crawlStatus, HTTP status, and latest event time | Classify 403/429/5xx responses for retry or manual investigation |
| Results retrieved | Retrieve all result pages after finished | Job totals, completed/error/skipped counts, raw result location, and checksums | Keep the job incomplete if counts do not reconcile |
| Knowledge published | Pass parsing, permission, deduplication, and indexing checks | Index version, page version, publication time, and rollback point | Retain the previous stable version |
This state machine prevents two common mistakes. First, jobStatus set to completed means execution has ended, not that every URL succeeded: the finish event may report both completed and errored counts. Second, HTTP 200 proves only a successful server response. It does not establish complete content, a correct canonical URL, matching language, parseable structured fields, or inclusion in the final retrieval index.
Record Market, Language, and Purpose in Job Configuration
Browser Run's /crawl endpoint supports source selection, depth, page limits, output formats, JavaScript execution, cache freshness, modification time, subdomains, and URL inclusion or exclusion rules. Avoid one unbounded job that mixes every language on a multi-market site. Register jobs by domain or language directory, and retain canonical URLs, hreflang, product or content entity IDs, and market versions together.
The endpoint also supports declaring crawlPurposes and contentUse to enforce the usage boundaries a target site expresses through Content Signals. Job configuration, event subscriptions, and result processing must maintain the same content purpose. Do not declare search or citation use during collection and then transfer the content into an incompatible training workflow. Record purpose-based denials as policy outcomes; do not bypass them by changing User-Agent or disabling restrictions.
The crawl source also affects how skipped counts should be interpreted. Cloudflare explains that a URL discovered and evaluated individually may be recorded as skipped when configuration rules exclude it. With a sitemap source, out-of-scope URLs may instead be omitted during bulk filtering. The skipped count is therefore not a complete inventory of uncovered URLs. Reconcile the target URL list or sitemap snapshot against the final result set.
A Minimal Event and Results Log
| Field group | Suggested fields | Purpose |
|---|---|---|
| Job identity | job ID, account, subscription, environment, brand, and target domain | Separate account-level events and preserve traceability |
| Content scope | Starting URL, source, limit, depth, include/exclude, and subdomain rules | Explain coverage and omissions |
| Content purpose | crawlPurposes, contentUse, robots rules, and access-policy version | Show that collection purposes match site signals |
| Page status | URL, crawlStatus, HTTP status, event time, and retry count | Investigate WAF rules, rate limits, broken pages, or configuration exclusions by error type |
| Result evidence | Totals, completed/error/skipped counts, cursor-completion flag, and raw result checksums | Prevent first-page-only retrieval or treating partial success as complete success |
| Knowledge publication | Parser version, index version, page version, publication time, and rollback point | Distinguish 'fetched' from 'searchable' |
Classify alerts by problem type. For a single 403 or 429 response, check authorization, WAF rules, and request rates. For repeated 5xx responses, check origin health. For many skipped URLs, review inclusion and exclusion rules. For limits or timeouts, review account limits, page counts, and job-splitting strategies. Switch downstream AI applications to a new version only when results are complete, error thresholds are met, all critical pages succeed, and knowledge-index checks pass.
How to Measure Whether the Change Works
Measure event subscriptions primarily through operational reliability, not AI citation rates. Track job-creation-to-started latency, page-change-to-updated latency, job duration, per-page success rates, error and skip ratios, result retrieval completeness, duplicate-event deduplication rates, knowledge publication latency, and rollback counts. These show whether the update pipeline is timely, complete, and recoverable.
Assess AI performance separately. After publishing a knowledge version, use a fixed question set to check retrieval hits, factual consistency, source links, and time-sensitive fields. Record your own application's metrics separately from citations and source visibility on Google, Bing, ChatGPT, Perplexity, and other public platforms. A successful Browser Run job proves at most that your collection process worked; it is not public AI search exposure.
Business Implications
This update moves web collection for internal AI knowledge bases from repeated completion polling to event-driven operations. For brands spanning markets and languages, with frequently changing product pages, per-page status and job summaries can expose 403 and 429 responses, broken URLs, accidental exclusions, and partial failures earlier. Content, engineering, data, and regional teams can coordinate around the same job ID. However, events are account-level, Queues defaults to at-least-once delivery, and events contain no page content. Without job registration, idempotent consumers, paginated result retrieval, and indexing checks, real-time notifications merely accelerate duplicate actions or the false conclusion that a finished job means updated knowledge. The value comes from linking events, results, page versions, and knowledge versions into an evidence trail.
Zhihe Growth's Assessment
The important change is not simply fewer polling calls. It is the ability to make AI knowledge updates observable, alertable, and replayable in production. Establish job identity, market and language scope, content purpose, and failure classification before connecting automatic indexing. Otherwise, faster automation only moves incorrect content into AI answers faster. GEO teams must distinguish three forms of evidence: Browser Run events establish the state of an internal collection job; complete results establish that the internal system retrieved page content; fixed-question tests and platform data establish whether content was correctly retrieved, presented, or cited in a particular AI experience. None substitutes for the others, and Cloudflare crawl results cannot establish access or ingestion by third-party platforms. This article can confirm the event types, account-level subscription scope, field structures, at-least-once delivery, result retrieval method, and job limits published by Cloudflare as of September 28, 2026. It cannot confirm each account's actual latency, cost, event ordering, third-party crawl status, or business conversion results.
Recommended Actions
- Register the job ID, target domain, environment, market, language, responsible owner, and configuration hash for every production crawl.
- Subscribe to crawl.started, crawl.updated, and crawl.finished. Explicitly separate test and production jobs in the consumer.
- Build idempotency keys from the job ID, event type, URL, status, and time fields so duplicate deliveries cannot cause duplicate ingestion or publication.
- Save the full crawlConfig when started arrives, and set maximum runtime and operational alert thresholds.
- Classify updated events by 403, 404, 429, 5xx, disallowed, and skipped. Do not hide policy errors behind blanket retries.
- On finished, mark the job only as awaiting retrieval. Fetch complete results by job ID and process every cursor page.
- Reconcile the target URL list or sitemap snapshot against the results. Verify critical product, service, evidence, and policy pages separately.
- Retain crawlPurposes, contentUse, robots, WAF, and content-access policy versions to prevent drift in purpose or access rules.
- Switch the production knowledge base only after parsing, permissions, deduplication, factual consistency, and indexing tests all pass. Keep a rollback point.
- Report collection success, knowledge publication latency, internal AI answer quality, and public-platform citation performance separately, not as one 'GEO growth' figure.
Limitations
Browser Run's `/crawl` endpoint is currently labeled Beta. Product fields, limits, and event schemas may change; consumers should inspect eventSchemaVersion and handle unknown versions conservatively. Event subscriptions are a push-based status channel, but Queues defaults to at-least-once delivery and occasional duplicates are possible. Do not assume each event arrives only once. Lifecycle events contain no crawled page content. After a job ends, retrieve complete results by job ID within the retention window. Current documentation allows a maximum runtime of 7 days and retains results for 14 days after completion. Pagination, errors, skips, and configuration filtering all separate 'job finished' from 'all target content retrieved.' This article concerns companies using Cloudflare Browser Run for their own search, knowledge-base, or RAG update workflows. It does not establish that Cloudflare AI Search, Google, Bing, OpenAI, or any third-party system has indexed the same content, nor does it guarantee rankings, AI citations, recommendations, clicks, inquiries, or revenue. Private, paid, personal, or regulated content also requires separate authorization, access-control, retention, and regional-compliance assessments.