Test four language directions
Chinese queries against Chinese records, Chinese against English, English against Chinese and English against English are separate configurations. Testing one alone cannot separate query-language effects from corpus-language effects. Translations must preserve model identifiers and constraints.
The dataset contains twelve synthetic product facts and twelve bilingual questions. Ten questions have author-defined relevant records; two have no answer in the corpus. Fields include voltage, capacity, connectors, materials, order quantity and weight. Four configurations produce forty-eight rankings, not forty-eight independent products or external-platform runs.
Record the model and execution conditions
The executed model is sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2, with a fixed revision in model-revisions.json. The program normalizes embeddings and ranks by cosine similarity, retaining the first three records and scores.
No ChatGPT, Google AI or Perplexity calls, answer generation or reranking were performed. This is a public model, not a Zhihe Growth foundation model. Reproduction requires the same revision, text and scoring method. Download time must not be presented as pure inference latency.
Actual results and denominators
| Query language | Corpus language | Top-1 hits on answerable questions | Recall@3 |
|---|---|---|---|
| Chinese | Chinese | 8/10 | 10/10 |
| Chinese | English | 7/10 | 10/10 |
| English | Chinese | 8/10 | 10/10 |
| English | English | 8/10 | 10/10 |
The two unanswerable questions are excluded from these denominators but retain their rankings. A searcher always returns similar records; similarity does not answer missing warranty or CE-certification questions. No abstention threshold was trained, so a high score must not invent absent facts.
Inspect failures rather than only the average
Q4 asks for a J connector; some configurations rank a neighboring model's connector record first. Q6 asks about dry-only suitability and also produces errors. Q5, concerning a restriction on untreated wood, misses top one in Chinese-to-English and English-to-English configurations. Inspect original text and scores rather than concluding that all negative statements fail.
Perfect Recall@3 on this small answerable subset does not guarantee a correct final answer. A generator could favor the first result, mix models or drop a restriction. That requires a separate generation-and-support experiment, which was not run here.
Improve retrieval without contaminating evaluation
First verify identifiers and constraints. Then consider exact matching, bilingual aliases, hybrid retrieval, reranking or structured filtering. Do not repeatedly rewrite these twelve test questions after inspecting failures and report the resulting fit as general improvement. Separate development examples from a held-out evaluation set.
Check translations as well: capacity, net weight and packaged weight are distinct facts. “Connector J” must not become merely “compatible connector.” Report directions, intents and unanswered questions separately so an aggregate does not hide a market-specific weakness.
How Zhihe Growth can apply the method
Zhihe Growth can use such experiments to diagnose whether English evidence maps correctly to Chinese purchasing questions or whether Chinese facts support equivalent English assets. Client delivery requires authorized facts, representative questions and independent review, with appropriate access controls.
Export-oriented GEO needs verifiable English evidence, not translated titles alone. Combine terminology control, hybrid retrieval and unit validation to locate the failing layer. Local ranking improvements still do not guarantee external citations.
Reproduction checklist
| Item | Correct practice | Distortion |
|---|---|---|
| Corpus | Twelve fact identifiers | Count translations as products |
| Denominator | Ten answerable, two separate | Delete unanswerable rankings |
| Directions | Report all four | Report only the best |
| Relevance | Fixed labels and revision | Relabel after seeing rankings |
| Truncation | Consistent model settings | Omit changed truncation |
| Output | Top three with scores | Keep only a hit flag |
Inspect each failure's question, relevant record and higher-ranked neighbor. Model numbers, voltage and negation may suggest exact constraint checks before a larger model. Preserve the baseline and regressions when testing improvements rather than overwriting inconvenient results.
Results and reference
Before comparing runs, verify data checksums and model revisions. A change to either requires a new experiment version; do not overwrite results and retain conclusions from the previous report.
Download the executed package. model-results.json contains forty-eight rankings; benchmark.json contains records and author labels. See Sentence Transformers documentation. The numbers describe this fixed teaching set, not population-level performance.