Hybrid retrieval combines candidates; it does not guarantee that keywords plus vectors always outperform either approach. Lexical search may preserve an exact identifier, while dense retrieval finds a differently worded explanation. Both can still omit a critical restriction. This experiment compares routes against one fixed relevance key and retains failures.
Why raw scores are not interchangeable
SQLite BM25 is lower-is-better, while normalized embedding dot products are higher-is-closer. Reversing one direction does not align their distributions. Adding the values can let the larger numeric scale dominate. Score normalization is possible, but needs validation data and distribution monitoring, not one favorable example.
Reciprocal rank fusion uses positions rather than raw scores. A document receives a contribution from each ranking in which it appears. Azure's hybrid ranking documentation explains the mechanism; it does not establish how unrelated commercial AI systems rank sources.
Work through a small example
lexical: A, B, C
dense: B, D, A
RRF(d) = sum(1 / (c + rank_i(d)))
c = 60; ranks start at 1
A receives 1/61 + 1/63, B receives 1/62 + 1/61, and D receives only 1/62. B benefits from appearing near the top of both rankings. This is agreement between retrieval routes, not proof of electrical compatibility or certification. The constant 60 is a disclosed teaching setting, not a universal optimum.
A duplicated document within one list must not receive multiple votes. The sample uses its first original rank and skips later duplicates without compressing subsequent ranks. Ties use document IDs for deterministic results. Without these choices, rerunning a small experiment can produce misleading differences.
What enters the actual fusion window
The program takes five lexical and five dense candidates, then sends the first five fused candidates to the reranker. The collection has only nine records. This is a teaching window, not a recommended production size. Dense retrieval scores all nine; lexical search may return fewer than five. It does not invent zero-score lexical matches to fill the list.
| Output | Meaning | Incorrect interpretation |
|---|---|---|
| bm25 | Lexical order | Factual accuracy |
| dense_scores | Similarity for each document | Probability that an answer is trustworthy |
| rrf | Fusion of two top-five lists | A ranking of every possible source |
| reranked | Reordered fused top-five list | Recovery of missing candidates |
Inspect the inputs and labels alongside the execution record. Q10 lacks an answerable source but still receives fused candidates. A ranking can exist when a supported answer cannot.
Decide whether fusion is worthwhile
Compare Recall@3, MRR@3 and nDCG@3 on the same questions. Then inspect identifier, condition, semantic and unknown-answer groups. An improved average does not justify deployment if the electrical-constraint group becomes worse. This small sample cannot support a marketing claim about increased public citation rates.
Production measurements should separate corpus encoding, index construction, retrieval, fusion and reranking. This package's lexical time includes index construction and cannot be compared directly with a dense query against precomputed vectors. The single-run CPU observations demonstrate execution, not stable service performance. A production comparison needs separate timers, warm-up, repeated runs, latency percentiles and controlled threading and caches.
Query-level gains and regressions
For Q2, lexical top-three results D1, D2 and D8 cover only one relevant source. Fusion returns D1, D9 and D2, raising Recall@3 from 0.5 to 1 in this synthetic setting. It is not a 50-percentage-point gain in public website citations.
Q1 shows a regression: dense results D9, D2 and D1 cover both labeled sources, but fusion returns D9, D2 and D3. Agreement on distractors can displace useful evidence. Fusion is not an oracle that selects the better route per query. Q9's Chinese stress query also remains unresolved at the top-three level.
Reporting only Q2 would be selective, so all ten queries remain downloadable, including no-answer Q10. Choose fusion constants and windows on independent validation data, then evaluate held-out queries. Repeatedly tuning on these ten questions would fit the teaching set, not establish generalization. Production work additionally needs larger catalogs, update-latency checks, duplicate-version handling and access constraints.
Audit duplicates and missing ranks
Deduplicate by stable document ID, not a shortened title. A document present in both lists receives both contributions; a duplicate within one list contributes only at its first rank. An absent document contributes zero, not an invented sixth place. Identical titles do not establish that two manual versions are interchangeable.
Print the original lists, then each contribution and total, then the final order and cutoff. Publish a deterministic tie rule; this lab uses document ID. Record whether filtering happens before or after fusion: the former changes candidate supply, while the latter can leave fewer than k results. A smaller constant c changes the influence of leading ranks, not necessarily accuracy. Separate tuning questions from held-out evaluation questions. This run fixes c at 60 rather than selecting a favorable configuration on the ten examples.
Zhihe Growth treats precise product facts and buyer-scenario language as complementary content responsibilities. One supports identity checks; the other makes relevant evidence discoverable through different questions. A website need not expose internal ranking weights, and an RRF score is not customer value. For Chinese exporters, model-to-manual relationships, factual scope and verifiable evidence come before algorithm selection.
Read about reranking boundaries, the RAG evidence path and retrieval metrics. No result here proves that an external platform will preferentially cite Zhihe Growth because of a particular page format.