Dense retrieval can connect related expressions without exact word overlap. It is not a product identity verifier. This guide provides actual nearest-neighbor outputs from an open-source sentence encoder on synthetic data. Manually assigned scores are not presented as embeddings, and no commercial AI citation test is included.

What encoding and cosine similarity compare

An encoder maps text into a vector. When vectors are normalized to unit length, their dot product equals cosine similarity. The score measures proximity in that representation space. It is not the probability that a specification is true, and it does not enforce a purchasing contract's constraints.

vectors = encoder.encode(texts, normalize_embeddings=True)
query_vector = encoder.encode(question, normalize_embeddings=True)
scores = vectors @ query_vector

The lab uses sentence-transformers/all-MiniLM-L6-v2, with its exact revision in the model manifest. See the model card and Sentence Transformers documentation for the encoding interface. Re-encode the collection when changing models. Vectors from unrelated models are not interchangeable coordinates.

Identifiers, units and negation make useful stress cases

AX-220, AX-110 and AX-220B share substantial vocabulary. A nearest-neighbor list may capture that all concern cleaners without treating supply voltage and battery voltage as non-substitutable facts. A passage saying “not for outdoor use” is evidence of a restriction, not support for an outdoor recommendation.

Stress case Required check Unsafe shortcut
Similar model IDs Original identifier, product ID, manual relationship Delete suffixes and merge products
Same unit, different object Separate supply, charger and motor fields Put every V value in one field
Negated condition Preserve the subject and polarity Match only “outdoor” or “cracked paint”
Cross-language query Check language coverage and preserved numbers Treat one English model as all Chinese retrieval

These are evaluation risks, not a claim that every embedding model fails them. Findings must identify the particular model revision and inputs. A different model, corpus or wording can change the result.

Reproduce a fixed evaluation without favorable rewording

The dataset contains nine English records and ten fixed queries. The program retains the complete nearest-neighbor list and scores in the run record. Freeze relevance labels before examining model results. Do not relabel inconvenient results merely to improve a metric.

Q9 is a Chinese stress query, not a multilingual benchmark. Q10 asks for warranty and price information absent from the corpus. Dense retrieval still returns a nearest document. Nearest does not mean relevant enough, and relevant does not mean sufficient to answer. This lab supplies no production abstention threshold: that requires calibration against an independent set containing genuinely unanswerable questions.

Put a fact gate after candidate retrieval

Separate identity and scope from semantic ranking. Identify the requested model and market, retrieve evidence, then check whether the required fields are supported. If a suffix is ambiguous, clarify the question or present a conditional comparison. Keep rejected candidates and rejection reasons so that a strict filter does not silently discard a valid manual.

Raising a universal similarity threshold is not a complete hallucination defense. An almost identical but incompatible product can score highly, while a terse table containing the exact answer can score lower. Verify the object, conditions, source version and claim support. The SSOT validation guide shows how to keep these checks distinct from retrieval scores.

Observed neighbors: the wrong bearing comes first

Q5 asks for BR-6204 bore diameter. Dense top-three results are D8, D7 and D6. D8 describes BR-6205 at 25 mm; correct D7 describes BR-6204 at 20 mm. Selecting the first result as the answer would misstate this synthetic diameter by 5 mm, despite perfect top-three recall. This is not real bearing compatibility advice.

For Chinese Q9, dense top-three results are D2, D1 and D3, excluding D9. Language, similar identifiers and electrical conditions vary together. Isolate these factors before attributing the error to language alone. One question cannot establish that all English encoders fail Chinese queries.

Retain the question, candidate IDs and actual fields rather than only an average similarity. A wrong first result with useful later evidence suggests testing reranking or constraints; absent evidence suggests indexing or recall work. Changes to preprocessing, normalization, model revision or candidate size require new records, not overwritten outputs. The published interpretation concerns this run, not guaranteed stability.

Zhihe Growth's application and limits

Set an error policy before deployment. Confusing power supply, load limits or compatible parts warrants conservative handling; a general introductory question may tolerate a broader candidate set. One similarity threshold across every task cannot establish that a product fact has been verified.

For Chinese companies expanding internationally, Zhihe Growth must make relevant scenarios discoverable while preserving model IDs, units and restrictions from Chinese source materials. The experiment provides inspectable failure cases for that editorial work. It is not evidence of a proprietary foundation model or a reconstruction of an external platform.

A more demanding evaluation would use authorized real materials, independent annotators, multilingual encoders and more confusing near-identical products. This small lab is educational, not proof of general industry gains. Compare lexical retrieval, rank fusion and reranking before deciding whether the problem lies in candidate retrieval or evidence judgment.

Knowledge center · GEO services · Research and evidence