Choose the unit before calculating a score
Official-domain citation rate counts answers that cite the domain. Citation precision asks whether supplied references support their associated statements. Fact coverage asks whether essential statements have sufficient evidence. A response can perform well on the first measure and fail both of the others.
Split independently checkable assertions. “AX220 has a 20 L tank” and “AX220 can operate outdoors” should not be a single claim. Create an edge between each claim and each cited source, labeled supported, partial, conflicting, irrelevant or unknown. One page can support multiple claims; repeating the same page against the same claim creates no additional edge.
State both denominators explicitly
The downloadable implementation uses a deliberately strict convention. Precision is fully supported unique claim–source edges divided by all unique citation edges. Coverage is required claims with at least one fully supporting edge divided by all predefined required claims. Partial support is not full support.
Unknown edges remain in the precision denominator and receive a separate count. This prevents inaccessible sources from disappearing and improving the apparent result. It is a conservative convention for this lesson, not a claim that every research benchmark uses the same method. With no edges, precision is null rather than zero or perfect; coverage is null when no required claims were specified. Compare runs only when claim segmentation, required facts and unknown handling match.
Recalculate the four-claim fixture
The execution record contains synthetic labels, not customer outcomes.
| Required claim | Source | Label | Strict treatment |
|---|---|---|---|
| c1 | s1 appears twice | Supported | One supported edge |
| c2 | s2 | Conflicting | One unsupported edge |
| c3 | s3 | Unknown | One unknown edge |
| c4 | None | Missing | Required claim without an edge |
After deduplication there are three edges, of which one is supported: precision is 1/3. Only one of four required claims is supported: coverage is 1/4. Missing claim c4 affects coverage but not the number of supplied edges. Counting links alone would reward the duplicated s1 and conceal this distinction.
Partial and joint support require additional care
A source proving capacity does not establish outdoor suitability. Split the assertion where possible; otherwise retain a partial label and explanation. Do not silently convert partial evidence into full support.
Joint support can require several documents, such as one identifying a certificate holder and another recording validity. The teaching function implements single-edge full support, not evidence-set reasoning. A joint-support workflow needs group identifiers, required documents, reasoning steps and reviewer decisions. Conflicting sources also require scope and version checks; a majority of copied pages does not settle the fact.
Separate semantic review from arithmetic
Software can deduplicate edges and calculate fixed formulas. A URL alone cannot tell it whether the document supports a claim. Retain the exact assertion, source title, resolved URL, supporting passage or page, applicable period, label, reasoning and review state. Resolve title-only source cards by opening or inspecting their links. An unresolved link remains pending, not absent.
A production evaluation benefits from independently assigned labels and recorded adjudication. The public example has author-defined labels and does not claim independent human agreement. Model-assisted review also needs bias checks described in the judge-bias guide. Preserve complete commercial-platform answers separately; this synthetic fixture is not part of a brand citation-rate denominator.
How Zhihe Growth uses these distinctions
Zhihe Growth separates official-domain citation, factual correctness and coverage of required evidence. For exporters, model identifiers, certification scope, units and service limitations often deserve priority, but the required facts must be defined before inspecting results.
Low precision suggests weak claim–source matching. Low coverage suggests missing facts or evidence. If both deteriorate, first check segmentation and document versions. Repeat the same question protocol after changes instead of claiming a causal uplift from one page. The support-audit guide covers claim review; the confidence-interval guide covers uncertainty.
Review exercise: identical scores can hide different problems
| Finding | Precision | Coverage | Next action |
|---|---|---|---|
| Repeated source edge | Unchanged after deduplication | Unchanged | Retain locations, deduplicate scoring |
| Required fact without source | No new supplied edge | May fall | Supply evidence or revise claim |
| Citation supports another model | Not supported | No support for this fact | Correct entity matching |
| Inaccessible source | Unknown remains in denominator | Not supported | Recover or review |
| One page supports two claims | Review two edges | May cover both | Do not count URLs instead of edges |
Record severity separately. A thousandfold unit error and an unsupported background statement may each count once but have different purchasing consequences. Disclose weighting rather than silently changing it. A changed required-fact list needs a new scoring version across the entire run.
Reproduce and interpret responsibly
Download the material package and run controls.py. It demonstrates accounting rules, not an autonomous truth-verification product connected to customer data.
The ALCE paper motivates evaluating cited generation along distinct dimensions. Any adapted metric still needs its own labeling and scoring specification. These results establish neither external-platform ranking nor commercial impact or causality.