Recall@3 of 1 does not mean a website has a 100% citation rate on public AI platforms. It means all labeled relevant records were retrieved within three positions for a particular corpus and key. Renaming an engineering measure as a marketing outcome destroys comparability.

Three ranking measures answer different questions

Suppose A has relevance grade 2, B grade 1, and all other records grade 0. For ranking X, A, B at k=2, only A is retrieved within the cutoff. Recall@2 is 1/2 and reciprocal rank is 1/2. Their equality is incidental: one measures coverage and the other the first relevant position.

Recall@k = retrieved relevant items / all labeled relevant items
RR@k = 1 / rank of first relevant item; none -> 0
DCG@k = sum((2^grade - 1) / log2(rank + 1))
nDCG@k = DCG@k / ideal DCG@k

Here DCG is 3/log2(3), ideal DCG is 3 + 1/log2(3), and nDCG is approximately 0.521. MRR averages reciprocal ranks across queries. The program's per-query mrr field is that query's contribution to a macro-average, not a complete benchmark MRR by itself.

No relevant label does not mean a perfect answer

Without any relevant records, Recall and nDCG have undefined denominators. The lab records null values and keeps the no-answer query separate instead of assigning a score of 1. Abstention needs a separate evaluation task; every query need not have a best answer.

Repeated IDs cannot increase recall. The metric function rejects duplicate document IDs, while fusion deduplicates within each input ranking. The labeling policy must also say whether negative evidence counts: a manual establishing that a product is unsuitable can be highly relevant to a compatibility question.

Use explicit names for website GEO measures

Measure Event and denominator used here Not equivalent to
Brand mention rate Valid answers mentioning the brand / all valid answers Official links
Official-site citation rate Valid answers with a confirmed official-domain citation / all valid answers Correct citation content
Target-page citation rate Valid answers citing the specified page / all valid answers Share within citing answers
Conditional target-page share Answers citing the target page / answers citing the official site Overall citation rate
Source-review coverage Reviewed source records / records needing review Support correctness
Claim support proportion Fully supported reviewable claims / all reviewable claims Absolute source truth

This is a proposed harmonized dictionary, not permission to rewrite historical denominators. Preserve older definitions and explain differences. Standardize future measurements explicitly. A published percentage alone cannot recover its original sample size.

Keep missingness visible

A normal answer with no official citation can be zero. A failed request with no answer is a technical failure. A visible source card whose URL is unresolved remains pending. Do not turn pending into zero or silently remove failures so that only convenient observations remain.

With N valid answers, K confirmed official citations and U unresolved source records, report a current lower bound K/N and upper bound (K+U)/N. State technical failures separately. These are missing-data bounds, not confidence intervals. Statistical uncertainty is addressed in the interval guide.

Check definitions against actual lab results

In the execution record, Q5 asks for BR-6204 bore diameter. Dense retrieval returns D8 before the correct D7. Recall@3 is still 1, but reciprocal rank is 0.5. A correct document somewhere in the first three is not the same as a correct first result.

For Q1, dense top-three results include both D9 and D1, whereas fusion omits D1 from its top three. Fusion does not guarantee monotonic improvement. The bundle includes code, inputs and environment records. Nine synthetic records do not support commercial performance claims.

Decide weighting before aggregation

A macro average gives each of the nine labeled questions equal weight. Pooling retrieved relevant items over all relevant items weights questions by the size of their relevance set. Those answer different questions and should not share an ambiguous “average recall” label.

This package retains per-query values rather than selecting a flattering aggregate. Keep Q10 as an explicit no-answer stress case. Do not convert its null scores to zero or erase it from the sample description. Ranking quality can be averaged over nine labeled questions, but abstention accuracy requires an answer stage that was not executed.

Citation aggregation has similar choices. Pooling platforms weights those with more answers; averaging platform percentages weights small and large samples equally. Predefine business weights and retain platform-level figures. Separate branded from unbranded questions so adding easy brand questions cannot masquerade as improvement. An acceptance register should specify objective, observation unit, numerator, denominator, missing-data rule and version. Matching metric names are not enough for comparability.

Zhihe Growth's decision framework

Zhihe Growth should keep diagnostic engineering measures separate from client acceptance measures. The former localize source and retrieval problems; the latter require real platform observations. Qualified inquiries need independent traffic, form and CRM records, not inference from a single citation.

See the Stanford evaluation chapter and BEIR, then continue with support auditing and before/after attribution. No new public-platform citation rate is reported here.

Knowledge center · GEO services · Research and evidence