Retrieval systems often fetch candidates and then apply a more detailed ranking model. Reranking asks which supplied candidates best match a question. It does not establish that specifications are true or ensure that a generated answer displays citations. This guide uses an actual open-source model on synthetic records, with inspectable configuration and outputs.
Bi-encoders and cross-encoders do different work
A bi-encoder represents queries and documents separately, allowing corpus vectors to be cached. A cross-encoder processes a query together with a candidate document. It can compare relevance within a bounded candidate set, but incurs per-pair computation. These are different workloads rather than interchangeable speed measurements.
The experiment uses cross-encoder/ms-marco-MiniLM-L6-v2; consult its model card. Its scores rank this set of candidates. They should not be presented to buyers as calibrated probabilities of truth. The model's training task differs from a client's technical catalog, so domain performance needs separate testing.
Compare the same candidates
Keep the first five fused results and rerank those same five. If a baseline receives three documents while the treatment retrieves twenty, the observed difference combines coverage and reranking. It cannot be assigned entirely to the cross-encoder.
candidates = fused_ranking[:5]
pairs = [(question, documents[doc_id]) for doc_id in candidates]
scores = reranker.predict(pairs)
The reordered result remains a subset of those five IDs. If D9 is sixth in the fusion output, the reranker cannot place it first because it never receives D9. Diagnose candidate coverage first, ranking second, and preservation of conditions in the final answer third.
Inspect the full execution record
The program, dataset and labels, model revisions and outputs are available together in the download bundle.
| Check | Question | Failure to retain |
|---|---|---|
| Coverage | Did necessary evidence enter the top five? | Correct source outside the window |
| Ordering | Did useful existing candidates move up? | A confusing product moved up instead |
| Top-three metrics | Did Recall, MRR and nDCG improve together? | Improvement in only one metric |
| Unknown answers | Are high-scoring candidates still returned? | Q10 has no warranty or price evidence |
| Latency | What computation did reranking add? | Reporting only the fastest observation |
The lab does not generate a natural-language answer. Ranking changes cannot be described as improved answer accuracy. No new measurement of public-platform website citations is included either.
Fact verification remains necessary
Q2 asks whether AX-220 can use 110 V. An AX-110 record may look relevant because it mentions 110 V, but it is not a substitute for the AX-220 manual. Lock the model, object, negation and version separately from relevance.
Q7 asks which hook-rod set includes a glue gun. D6 provides evidence that one is not included; it is not a recommendation satisfying the requirement. A relevance key can legitimately include evidence supporting refusal or a restriction. State that annotation rule in advance so reviewers do not treat useful negative evidence as irrelevant.
Record cost without false precision
The CPU timings describe one run and include machine and package versions. They are not throughput, network-service, concurrency or price measurements. Do not convert milliseconds into unmeasured cloud costs. A production test needs fixed candidate counts, warm-up, repeated measurements, median and tail latency, and memory and maintenance costs.
Model inference is only part of the expense. Correct labels, conflicting manuals and restrictions need ongoing review. For a small catalog, repairing a missing specification table may be more valuable than adding a reranking service.
Observed changes include a regression
Q1 moves from fused D9, D2, D3 to reranked D9, D3, D1. Useful D1 returns to the first three, but distractor D3 remains ahead of it. Better coverage does not establish safe product recommendations.
Q2 moves from D1, D9, D2 to D1, D2, D9. Recall stays unchanged because both relevant records remain present, but nDCG falls from 1 to approximately 0.920. A reranker can worsen an already ideal ordering. Q9 recovers D9 within three positions but still starts with D2, so the Chinese compatibility problem is not fully resolved.
Q10 still receives candidates despite missing warranty and price evidence. Production use requires a separate sufficiency or abstention decision after ranking. These observations concern pinned models and the recorded CPU environment, not every cross-encoder or a comparison between GEO vendors. Expand evaluation with independent questions and review rather than deleting inconvenient failures.
What this means for Zhihe Growth
Acceptance should not depend on average rank alone. Check identity mismatches, lost negation and insufficient candidates separately, and decide whether each blocks a recommendation. Sensitive parameters still require approved-fact checks. An uncalibrated ranking score must not be displayed as a probability of correctness.
Zhihe Growth publishes this as engineering teaching material for exporters diagnosing retrieval, ordering and evidence-support failures. Deploying an open-source model is not described as inventing a proprietary foundation model. A production reranker should be justified by an evaluation on authorized materials.
Continue with hybrid retrieval, citation support and metric definitions. When an external platform does not disclose its reranker, its implementation remains unknown.