Skip to main content
Industry Insights | Zhihe Growth Research Center

Cloudflare AI Search's 64-Byte String Filter Prefix: Designing Multilingual Metadata

The source article describes a shared 10 KiB metadata envelope per vector and filtering on only the first 64 UTF-8 bytes of string values. It covers field design, encoding, reindexing and validation.

Key Takeaway

The cited Cloudflare AI Search documentation allows up to 10 KiB of shared metadata per vector, but only the first 64 UTF-8 bytes of each string are filterable. A stored value is therefore not necessarily filterable in full. Multilingual brands should use short, stable, controlled values for market, language, content type, product family and publication status, keep long descriptions in content or non-filter fields, and plan a full reindex and query regression tests after metadata schema changes.

Facts and background

Cloudflare's August 25, 2026 release notes describe a shared 10 KiB compact-JSON UTF-8 metadata envelope per vector. This is not a separate allowance for each field: system metadata, field names, JSON syntax and custom metadata share the space. Strings can be stored beyond their filterable prefix, but filter conditions can match only the first 64 UTF-8 bytes of each string.

These are separate limits: 10 KiB governs metadata storage per vector, while 64 bytes governs each string's indexed filter prefix. A long product-family name may be stored fully or partly but fail the intended filter if its market, version or status code occurs after byte 64. Bytes are not characters, particularly for Chinese, Japanese, Korean and emoji: UTF-8 character lengths vary.

Separate built-in metadata, custom fields and content

Information layerBehavior in the cited documentationAppropriate usesNot a substitute for
Built-in metadataAutomatically extracts filename,folderandtimestampFile identity, path groups and modification-time rangesCustom market, language or permission semantics
Custom metadataUp to five custom fields per instance, supporting text,number,booleananddatetimePre-retrieval conditions such as market, language, type, product family and publication statusLong explanations, evidence text or arbitrary tag collections
Page or document bodyCrawled, converted, chunked and indexed for retrievalComplete product facts, restrictions, cases, evidence and explanationsTenant or permission boundaries requiring strict exclusion before retrieval

Metadata filters narrow the candidate set before vector, keyword or hybrid retrieval. They can impose conditions such as searching only public English-language knowledge for Southeast Asia, but do not change the content's semantic relevance. Similarity thresholds, reranking, query rewriting and generation prompts remain separate settings. More metadata does not automatically improve answers.

Put stable short codes within the first 64 bytes

Filtering fields should prioritize stability, uniqueness and comparability over readability. Instead of setting market to a long description such as 'English content for Singapore, Malaysia and Indonesia', use controlled codes such as sea-en; keep human-readable descriptions in titles, content or the CMS. This reduces inconsistencies caused by translation, case, spaces, punctuation and excessive length.

Field names are case-insensitive and stored in lowercase; timestamp,folderandfilename are reserved names and cannot be used for custom fields. Count encoded UTF-8 bytes, not characters. If a long value is necessary, place the stable filter code first and store the display name separately; do not assume a keyword beyond the prefix remains filterable.

An illustrative five-field model for multilingual brands

FieldTypeExamplePurposeDesign constraint
markettextsg,de,globalRestrict the target marketUse stable market codes, not long regional descriptions
languagetexten,de,zh-cnRestrict the answer languageStandardize case and hyphens first
content_typetextproduct,evidence,faqDistinguish fact, evidence and Q&A pagesDo not turn temporary section names into permanent types
product_familytextpump-aRestrict retrieval to a stable product familyMap each code to a canonical product entity
is_publicbooleantrue,falseExclude unpublished contentNot an authentication or authorization system

This is an illustrative minimum, not a universal model. If time validity matters more than product family, replace a lower-value text field with a datetime field. For numeric product versions, consider number with range operators. The built-in timestamp already records the object's last modification time; reuse it where appropriate rather than spending one of five custom fields on duplicate information.

Multi-brand, multi-tenant or restricted data must not rely solely on is_public=false filtering. Misconfiguration, omitted conditions or schema changes can broaden the candidate set. Sensitive data needs separate instances, Cloudflare Access, server-side authorization, least privilege and source isolation. Metadata filtering supports retrieval governance, not a security boundary.

Metadata entry points differ across websites, R2 and built-in storage

Website sources extract custom fields from <head> package provides <meta> tags using the nameorproperty attribute, after defining the instance schema. Field names match the schema case-insensitively; the content value is converted to the configured type. Boolean fields accept true/1/yes and false/0/no; other values are invalid and omitted.

R2 sources use S3-compatible x-amz-meta-* headers, while built-in storage attaches metadata during Items API uploads. All sources must follow the instance schema and vector metadata limits. Inconsistent types or controlled values make filters unreliable; agree on a field dictionary before connecting multiple sources.

On web pages, title,descriptionandimage have recognized meta sources but still require schema definitions for extraction. Standard meta values take precedence over Open Graph equivalents. Prioritize display and summary fields separately from fields needed for pre-retrieval filtering, so duplicate, unstable or non-decision-making content does not exhaust the field allowance.

Schema changes have reindexing costs

Cloudflare documents that custom metadata schema changes trigger a full reindex: new fields enter the index, removed fields leave it, and existing vectors receive the new metadata structure. Do not experiment with field names, types or controlled values directly in a large, multilingual or frequently updated production instance.

Test extraction, type conversion, byte lengths and filters on representative samples before scheduling the change. Afterward, inspect actual Items metadata and test both inclusion and exclusion for each market, language, type and publication status. Finding desired content alone does not prove that incorrect-market or unpublished content is excluded.

Read filter operators carefully

The cited filtering syntax supports $eq,$ne,$in,$nin and range operators; assigning a field directly is equivalent to $eq. Multiple fields in one filter object are implicitly combined with AND.$in compares a stored scalar with a list of candidate scalars; it does not search inside a stored array. Cloudflare documents that Vectorize can store string arrays but does not index or filter those arrays.

Do not assume an untested string array can filter a page belonging to several product families. Consider clearer knowledge units, a primary family, multiple canonical pages or a different instance and path structure. Validate actual query results: valid-looking JSON does not prove filtering works.

An auditable release checklist

Check layerPass criteriaFailure symptomsAction
Field dictionaryAt most five fields, with defined names, types, values, owners and change rulesDuplicate synonymous fields, reserved-name conflicts or inconsistent typesUnify the schema before starting a full reindex
Capacity and prefixStable filter codes fit within 64 UTF-8 bytes; the shared envelope leaves room for system overheadLong suffixes fail to match or strings are truncatedShorten values, put stable codes first and reindex
Source extractionWebsite meta, R2 headers and Items API values enter the index as intendedInvalid booleans, missing fields or conflicting standard and Open Graph valuesCorrect source values and schema, then rerun sync
Filter regression testsExpected documents match; incorrect market, language, type and status samples are excludedOnly inclusion is testedAdd negative and boundary cases
Security and effectivenessAccess controls isolate restricted data; relevance and answer quality are measured separatelyTreating metadata as authorization or ranking optimizationSeparate security, retrieval and GEO metrics

Retain the field-schema version, source, sample URL or object, raw metadata, UTF-8 byte count, sync time, query filter, returned documents, expected inclusion or exclusion, and retest date. This distinguishes source-data, schema, sync, filtering and relevance problems after redesigns or changes in controlled values.

Separate filter correctness from GEO outcomes

Metadata filtering directly tests whether candidate documents satisfy conditions. It does not establish that Google, Bing, ChatGPT or another public system crawled, indexed or cited the page. AI Search is an enterprise-configured retrieval system. Exposing it through MCP or a public endpoint demonstrates behavior at that endpoint, not across public platforms.

First test extraction and exclusion, then retrieval relevance, returned sources and answer accuracy, and finally public brand mentions, citations, visits and conversions. Compare answers with the same questions and instance configuration before and after a change. Without a comparison, concurrent fluctuations cannot be attributed to a 64-byte prefix or metadata adjustment.

Impact on enterprises

First, the multilingual risk is often confusing 64 UTF-8 bytes with 64 characters, not exhausting 10 KiB. Chinese, Japanese, Korean, accented characters and emoji use different byte lengths; long localized descriptions make poor stable filter keys. Second, five custom fields are a limited governance resource. Prioritize dimensions that change the candidate set, such as market, language, content type, product family and publication status, rather than repeating content or built-in fields. Third, schema changes affect production reindexing. Treat renaming, type changes and field additions or removals as versioned changes with rollback, sync windows and regression tests. Fourth, filtering neither replaces access controls nor proves public AI-search visibility. Isolate restricted data at the source and access layers; verify GEO separately through public answers, source links, crawl and visit evidence, and site analytics.

Zhihe Growth's Assessment

The capacity update permits richer metadata, but reliable retrieval governance still depends on a small, stable field model within the five-field limit. For multilingual knowledge bases, do not aim to fill 10 KiB: use short codes for filtering, content for explanation, built-in fields for file identity and time, and access controls for security. Define each field's purpose, type, allowed values, source, owner and retirement policy. Test bytes and filters on representative pages before a full reindex. Once fields are shared across sites, R2 objects, uploads and query code, changes become costly; negative tests before launch matter more than extra descriptive fields. This article describes the cited AI Search metadata limits and indexing behavior. It does not establish that public search platforms read these fields or that configuration increases rankings, citation rates, brand mentions, traffic or conversions.

Recommended action

  1. Inventory market, language, content-type, product-family, version, time and publication-status fields; remove candidates that do not affect retrieval decisions.
  2. Use no more than five custom fields and document each field's type, allowed values, source, owner, default and retirement rule.
  3. Do not use the reserved names `timestamp`, `folder` or `filename` for custom fields. Reuse suitable built-in metadata first.
  4. Calculate UTF-8 byte lengths for text filter values. Keep stable codes within the first 64 bytes; do not rely solely on long localized descriptions.
  5. Standardize lowercase ASCII values for market, language, content type and product family. Keep synonyms, free text and temporary section names out of production filter values.
  6. Check field injection and type conversion separately for website meta tags, R2 `x-amz-meta-*` headers and Items API uploads.
  7. Save the schema version, instance configuration and baseline queries before changes; reserve a full-reindex window and rollback plan.
  8. Test positive, negative and boundary queries for each field combination, including documents that must be excluded.
  9. Protect sensitive, tenant-specific and unpublished data with separate instances, authentication, authorization and source isolation, not metadata filters alone.
  10. Record filter correctness, retrieval relevance, answer accuracy, public AI citations and site conversions separately; avoid cross-layer attribution.

Limitations

This article reflects the Cloudflare AI Search documentation cited as of September 4, 2026, when the product was described as Beta. Types, limits, APIs, pricing and reindexing behavior may change; check the release notes, Metadata attributes, Filtering, Website and Limits pages before implementation. The 10 KiB limit is a shared compact-JSON UTF-8 envelope per vector, including system metadata and JSON overhead, not net custom space or a separate allowance per document or field. Required system metadata takes priority, and configured strings may be truncated at UTF-8 character boundaries. Leave headroom rather than treating near-limit tests as reliable capacity. The 64-byte rule concerns each indexed string's filterable prefix, not whether the rest is stored or how other types behave. The cited documentation permits string-array storage but not array indexing or filtering. Metadata filtering limits AI Search candidates; it is not access control, a legal-compliance conclusion, public indexing or a citation guarantee.

Source Verification