Skip to main content
Industry Insights | Zhihe Growth Research Center

Cloudflare Adds Crawl Lifecycle Events: Building an Auditable AI Knowledge Update Pipeline

The September 28 source article describes Browser Run's crawl.started, crawl.updated, and crawl.finished events, with page-level status, failure queues, idempotent handling, and retained results for auditable knowledge-base updates.

Key Takeaway

Cloudflare now lets Browser Run `/crawl` jobs push start, per-page status update, and finish events to Queues, instead of relying only on periodic status polling. International brands that use web content for their own AI search, support knowledge bases, or RAG systems should use event subscriptions for scheduling, failure alerts, and auditing, then actively retrieve complete results by job ID after completion. The events contain no page content and do not prove that Google, Bing, ChatGPT, or other third-party platforms have crawled, indexed, or cited the brand's content.

Facts and Background

On September 25, 2026, Cloudflare announced that Browser Run /crawl jobs can publish lifecycle events to Cloudflare Queues. The three documented event types are crawl.started,crawl.updated and crawl.finished. They indicate that a job has started, a discovered URL has changed status, or the overall job has finished. Cloudflare positions subscriptions as a push-based alternative to status polling, for tracking progress or triggering downstream processing.

This changes a company's own web collection and knowledge-update workflow, not public search engines' indexing protocols. Events show collection progress, which URLs succeeded or failed, and when a job ended. They do not show whether a third-party AI platform has indexed content, how it understands it, or when it will cite it. Assess internal knowledge-base ingestion, public search access, and final citation performance separately.

What the Three Event Types Establish

EventKey fieldsSuitable actions to triggerWhat the event alone does not prove
crawl.startedjobId,createdAt,crawlConfigRegister the job version, lock the configuration, and start the timeout clockThat any page has been crawled successfully
crawl.updatedjobId,URL,crawlStatus,httpStatusUpdate per-page records, flag 403/404/5xx responses, and trigger targeted alertsThat page content is in the vector store or can support correct answers
crawl.finishedjobStatus, start and end times, total, completed, errored, skipped, crawlConfigReconcile totals, fetch complete results, and decide whether to proceed to parsing and indexingThat all target URLs are covered or the knowledge base has been published

Browser Run events are account-level: one subscription receives events for every crawl job in the account without a source-specific selector. A consumer cannot assume its queue serves only one domain, market, or knowledge base. At minimum, join the job ID to a registry containing the target domain, environment, language, market, content purpose, configuration version, and responsible owner. This keeps test, production, and different brands' data streams separate.

Event metadata also includes accountId, eventSubscriptionId, eventSchemaVersion, and eventTimestamp. Cloudflare Queues provides at-least-once delivery by default, so the same message may occasionally arrive more than once. Use idempotency keys and make duplicate events safe to replay. Two update deliveries do not mean a page was crawled twice and must not cause duplicate indexing or publication.

Events Are Not Content: Retrieve Complete Results After Completion

Cloudflare explicitly states that lifecycle events carry status information, not crawled page content. After receiving crawl.finished, use the job ID to call the results endpoint, retrieve every page of records, and handle URL statuses such as completed,errored,disallowed,skipped or cancelled separately. Responses exceeding 10 MB include a cursor. Do not mark content as fully ingested until all result pages have been retrieved.

Current documentation allows crawl jobs to run for up to 7 days and retains results for 14 days after completion. Saving events without retrieving and archiving the results within that window risks losing page content and per-page diagnostic evidence. A safer workflow moves a finished job only to 'awaiting retrieval.' Mark the knowledge base as updated only after complete retrieval, validation, parsing, deduplication, permission filtering, and indexing succeed.

Use a Five-Stage State Machine, Not One Completion Flag

StageEntry conditionRequired evidenceFailure action
Job registeredCreate the job and obtain its job IDTarget URLs, configuration hash, market, language, and initiation timeDo not enqueue downstream work
Crawl runningReceive started or query a running statuscrawlConfig, event time, and timeout deadlineAlert on timeout without assuming a cause
Per-page validationReceive updatedURL, crawlStatus, HTTP status, and latest event timeClassify 403/429/5xx responses for retry or manual investigation
Results retrievedRetrieve all result pages after finishedJob totals, completed/error/skipped counts, raw result location, and checksumsKeep the job incomplete if counts do not reconcile
Knowledge publishedPass parsing, permission, deduplication, and indexing checksIndex version, page version, publication time, and rollback pointRetain the previous stable version

This state machine prevents two common mistakes. First, jobStatus set to completed means execution has ended, not that every URL succeeded: the finish event may report both completed and errored counts. Second, HTTP 200 proves only a successful server response. It does not establish complete content, a correct canonical URL, matching language, parseable structured fields, or inclusion in the final retrieval index.

Record Market, Language, and Purpose in Job Configuration

Browser Run's /crawl endpoint supports source selection, depth, page limits, output formats, JavaScript execution, cache freshness, modification time, subdomains, and URL inclusion or exclusion rules. Avoid one unbounded job that mixes every language on a multi-market site. Register jobs by domain or language directory, and retain canonical URLs, hreflang, product or content entity IDs, and market versions together.

The endpoint also supports declaring crawlPurposes and contentUse to enforce the usage boundaries a target site expresses through Content Signals. Job configuration, event subscriptions, and result processing must maintain the same content purpose. Do not declare search or citation use during collection and then transfer the content into an incompatible training workflow. Record purpose-based denials as policy outcomes; do not bypass them by changing User-Agent or disabling restrictions.

The crawl source also affects how skipped counts should be interpreted. Cloudflare explains that a URL discovered and evaluated individually may be recorded as skipped when configuration rules exclude it. With a sitemap source, out-of-scope URLs may instead be omitted during bulk filtering. The skipped count is therefore not a complete inventory of uncovered URLs. Reconcile the target URL list or sitemap snapshot against the final result set.

A Minimal Event and Results Log

Field groupSuggested fieldsPurpose
Job identityjob ID, account, subscription, environment, brand, and target domainSeparate account-level events and preserve traceability
Content scopeStarting URL, source, limit, depth, include/exclude, and subdomain rulesExplain coverage and omissions
Content purposecrawlPurposes, contentUse, robots rules, and access-policy versionShow that collection purposes match site signals
Page statusURL, crawlStatus, HTTP status, event time, and retry countInvestigate WAF rules, rate limits, broken pages, or configuration exclusions by error type
Result evidenceTotals, completed/error/skipped counts, cursor-completion flag, and raw result checksumsPrevent first-page-only retrieval or treating partial success as complete success
Knowledge publicationParser version, index version, page version, publication time, and rollback pointDistinguish 'fetched' from 'searchable'

Classify alerts by problem type. For a single 403 or 429 response, check authorization, WAF rules, and request rates. For repeated 5xx responses, check origin health. For many skipped URLs, review inclusion and exclusion rules. For limits or timeouts, review account limits, page counts, and job-splitting strategies. Switch downstream AI applications to a new version only when results are complete, error thresholds are met, all critical pages succeed, and knowledge-index checks pass.

How to Measure Whether the Change Works

Measure event subscriptions primarily through operational reliability, not AI citation rates. Track job-creation-to-started latency, page-change-to-updated latency, job duration, per-page success rates, error and skip ratios, result retrieval completeness, duplicate-event deduplication rates, knowledge publication latency, and rollback counts. These show whether the update pipeline is timely, complete, and recoverable.

Assess AI performance separately. After publishing a knowledge version, use a fixed question set to check retrieval hits, factual consistency, source links, and time-sensitive fields. Record your own application's metrics separately from citations and source visibility on Google, Bing, ChatGPT, Perplexity, and other public platforms. A successful Browser Run job proves at most that your collection process worked; it is not public AI search exposure.

Business Implications

This update moves web collection for internal AI knowledge bases from repeated completion polling to event-driven operations. For brands spanning markets and languages, with frequently changing product pages, per-page status and job summaries can expose 403 and 429 responses, broken URLs, accidental exclusions, and partial failures earlier. Content, engineering, data, and regional teams can coordinate around the same job ID. However, events are account-level, Queues defaults to at-least-once delivery, and events contain no page content. Without job registration, idempotent consumers, paginated result retrieval, and indexing checks, real-time notifications merely accelerate duplicate actions or the false conclusion that a finished job means updated knowledge. The value comes from linking events, results, page versions, and knowledge versions into an evidence trail.

Zhihe Growth's Assessment

The important change is not simply fewer polling calls. It is the ability to make AI knowledge updates observable, alertable, and replayable in production. Establish job identity, market and language scope, content purpose, and failure classification before connecting automatic indexing. Otherwise, faster automation only moves incorrect content into AI answers faster. GEO teams must distinguish three forms of evidence: Browser Run events establish the state of an internal collection job; complete results establish that the internal system retrieved page content; fixed-question tests and platform data establish whether content was correctly retrieved, presented, or cited in a particular AI experience. None substitutes for the others, and Cloudflare crawl results cannot establish access or ingestion by third-party platforms. This article can confirm the event types, account-level subscription scope, field structures, at-least-once delivery, result retrieval method, and job limits published by Cloudflare as of September 28, 2026. It cannot confirm each account's actual latency, cost, event ordering, third-party crawl status, or business conversion results.

Recommended Actions

  1. Register the job ID, target domain, environment, market, language, responsible owner, and configuration hash for every production crawl.
  2. Subscribe to crawl.started, crawl.updated, and crawl.finished. Explicitly separate test and production jobs in the consumer.
  3. Build idempotency keys from the job ID, event type, URL, status, and time fields so duplicate deliveries cannot cause duplicate ingestion or publication.
  4. Save the full crawlConfig when started arrives, and set maximum runtime and operational alert thresholds.
  5. Classify updated events by 403, 404, 429, 5xx, disallowed, and skipped. Do not hide policy errors behind blanket retries.
  6. On finished, mark the job only as awaiting retrieval. Fetch complete results by job ID and process every cursor page.
  7. Reconcile the target URL list or sitemap snapshot against the results. Verify critical product, service, evidence, and policy pages separately.
  8. Retain crawlPurposes, contentUse, robots, WAF, and content-access policy versions to prevent drift in purpose or access rules.
  9. Switch the production knowledge base only after parsing, permissions, deduplication, factual consistency, and indexing tests all pass. Keep a rollback point.
  10. Report collection success, knowledge publication latency, internal AI answer quality, and public-platform citation performance separately, not as one 'GEO growth' figure.

Limitations

Browser Run's `/crawl` endpoint is currently labeled Beta. Product fields, limits, and event schemas may change; consumers should inspect eventSchemaVersion and handle unknown versions conservatively. Event subscriptions are a push-based status channel, but Queues defaults to at-least-once delivery and occasional duplicates are possible. Do not assume each event arrives only once. Lifecycle events contain no crawled page content. After a job ends, retrieve complete results by job ID within the retention window. Current documentation allows a maximum runtime of 7 days and retains results for 14 days after completion. Pagination, errors, skips, and configuration filtering all separate 'job finished' from 'all target content retrieved.' This article concerns companies using Cloudflare Browser Run for their own search, knowledge-base, or RAG update workflows. It does not establish that Cloudflare AI Search, Google, Bing, OpenAI, or any third-party system has indexed the same content, nor does it guarantee rankings, AI citations, recommendations, clicks, inquiries, or revenue. Private, paid, personal, or regulated content also requires separate authorization, access-control, retention, and regional-compliance assessments.

Sources