A useful product result must preserve the constraints that a buyer cannot change. All AX products in this article are synthetic teaching examples, not client products. The experiment explains controllable engineering decisions; it does not reproduce a commercial AI platform's ranking system.
What an inverted index and BM25 actually do
An inverted index maps terms to documents. BM25 ranks matching candidates using term frequency, rarity across the collection and document length. Repetition has diminishing returns; a longer document should not win simply because it contains more words. The Stanford information retrieval textbook provides the underlying retrieval concepts.
Inspect tokens before tuning parameters. A visibly intact model number is not necessarily indexed as one identifier. A tokenizer may split a hyphen. Putting the identifier in a title improves clarity, but it is not an equality constraint. Maintain the original model string in a dedicated field when exact identity matters.
Four records that invite the wrong conclusion
| Record | Model | Supported fact | Invalid inference |
|---|---|---|---|
| D1 | AX-220 | 220 V supply, 20 L tank, indoor sealed hard floors | Compatible with 110 V |
| D2 | AX-110 | 110 V supply, 20 L tank | Equivalent to D1 |
| D3 | AX-220B | 24 V battery, 220 V charger input | Charger input equals motor voltage |
| D9 | AX-220-manual | An appendix explaining AX-220 voltage restrictions | A separate product identity |
D9 exposes an important modeling problem. An exact product filter can preserve identity but exclude useful manuals if filenames are treated as model IDs. A production catalog needs a relationship between the manual and the product. The teaching code does not infer that relationship; implementing it is a useful extension exercise.
Establish a transparent SQLite baseline
Download the lab bundle, corpus and relevance labels, and program. The baseline uses SQLite FTS5, OR-connected query terms, model-field weight 4 and body-field weight 1. These are teaching choices, not optimized settings or a prescription for website keyword density.
CREATE VIRTUAL TABLE products USING fts5(id UNINDEXED, model, body);
SELECT id FROM products
WHERE products MATCH '"AX" OR "220"'
ORDER BY bm25(products, 0, 4, 1), id;
SQLite's BM25 function uses a lower-is-better convention. Its value is not a probability. Consult the implementation documentation before comparing scores from different systems. The program binds SQL parameters and constructs the search expression from extracted terms. Do not concatenate arbitrary customer input into SQL or an unrestricted search expression.
Calculate term-frequency saturation
One BM25 term-frequency factor has the form below; the complete score also incorporates inverse document frequency. SQLite returns a lower-is-better score, so it cannot be added directly to another system's positive score.
tf_factor = f * (k1 + 1) / (f + k1 * (1 - b + b * dl / avgdl))
f = 2, k1 = 1.2, dl = avgdl -> 4.4 / 3.2 = 1.375
f = 4, same lengths -> 8.8 / 5.2 = 1.6923
Doubling frequency from two to four does not double the factor. This illustrates saturation, not an instruction to repeat product names. The calculation omits IDF and is not a final query score from this package. Actual scores depend on corpus statistics, tokenization and field weights. A changed corpus can change scores without establishing that page quality changed; inspect the index version first.
Keep the full-question ranking, then run an identifier-only comparison. If an identifier query fails, inspect tokenization, case handling, aliases and indexed fields. If identity is correct but suitability is wrong, inspect the evidence and constraints. Adding more keywords does not repair a false electrical claim.
The call lexical_rank(documents, 'AX-220', exact_model=True) promotes records whose model field matches exactly. This is an explicit business rule in the program, not a semantic property of BM25. It still cannot invent warranty, price or certification information. Q10 asks about warranty and price, neither of which exists in the corpus. Ranking D1 first does not make those questions answerable.
Review failures as well as successful examples
The execution record contains rankings, query groups, top-three metrics and timings for every query. Both D1 and D9 support Q2's voltage restriction. Retrieving D1 alone does not recover all labeled evidence. Q10 has no relevant answer: its ranking metrics are null, not perfect, and the question remains visible in the record.
There are only nine English documents and ten authored queries, including one Chinese stress query. Lexical timing includes temporary index construction; dense timing covers query encoding and scoring. These are not comparable speed benchmarks. A real catalog evaluation needs a held-out set, varied identifier formats, independently checked labels, and a separate assessment of abstention on unanswerable questions.
An observed failure despite a model field
For Q1, “AX-220 supply voltage,” lexical top-three results are D9, D3 and D2; D1 is absent. OR-connected terms deliberately provide broad retrieval. Matching AX, 220, supply or voltage is not the same as belonging to the requested product. Model-field weighting does not implement identity constraints.
For Q5, lexical search puts correct D7 first while dense retrieval puts D8 first. Together these examples argue against choosing one favorable query to declare a universal winner. Compare the baseline, exact-model constraints, manual-to-product relations and ambiguity handling one change at a time. Preserve filtered evidence and use held-out questions before choosing weights. Hardcoding a bonus for AX-220 after seeing Q1 would contaminate evaluation.
The lexical implementation extracts ASCII letters and digits only, so Q9 does not represent production Chinese tokenization. On the website, the corresponding action is clearer model relationships and manual links, not repeating a model name ten times. Keep internal index changes separate from public-platform ranking claims.
How Zhihe Growth applies this distinction
For Chinese B2B exporters, Zhihe Growth's practical task is to align public model identifiers, versions, units, conditions and manual links. Retrieval experiments help distinguish missing facts from poorly separated facts. They do not imply control over Google, ChatGPT or Bing ranking weights.
Continue with embedding ambiguity, hybrid retrieval and retrieval metrics. The technical FAQ provides short answers. None of this lab's ranking scores is an official-website citation rate.