Define quality before asking for a winner

“Better” may mean correct facts, adequate citations, readable prose, brand prominence or suitability for a purchasing task. A single undifferentiated score can reward verbosity while concealing an incorrect specification.

Use separate dimensions for critical errors, citation support, required-question coverage, honest uncertainty and readability. Track brand presence separately, without increasing the factual score. A short answer that declines to invent certification can be preferable to an elaborate unsupported certification list. The judge's output is a recommendation; supporting records remain the authority.

Test three biases with controlled pairs

The LLM-as-a-Judge study discusses position, verbosity and self-preference risks. Its observed magnitudes should not be generalized to every current model, but these risks deserve local checks.

Present an answer pair in A/B and B/A order. Compare concise and expanded versions carrying identical facts. Replace brand names with neutral identifiers while preserving evidence. Change one factor at a time and save answer hashes, order, prompt, model version and outputs. An aggregate average after swapping can conceal question-level reversals, so keep the paired decisions.

Use a review worksheet without invented results

The package's review-templates.json defines fields for question, answer version, required facts, evidence passages, rubric, blinding map, presentation order, judge output, human labels and adjudication. Unperformed model or human assessments remain null, not illustrative scores later counted as observations.

Include clearly correct answers, deliberate factual errors, partial support, inaccessible sources, conflicting records and questions requiring abstention. Changing the synthetic AX220 voltage to 110 V creates an authored negative example, not evidence of an actual model failure. Software can validate paired records and score arithmetic; it cannot establish the reliability of a judge that was never run.

A gold standard requires an actual review process

Reviewers need the same definitions and frozen evidence, independent labeling, and explicit adjudication. Report sample size with raw agreement. Cohen's kappa can adjust for chance agreement, but small or highly imbalanced samples require caution. Agreement does not prove that the rubric itself is correct.

The public material includes two authored columns of six labels for teaching agreement calculations. They are not observations from two independent people. Real projects should retain private reviewer identifiers and decision histories. Certification, legal qualifications and safety-relevant specifications need authorized evidence or appropriately qualified review, not a model's final vote.

Respond to failed calibration

Frequent order reversals suggest narrowing the task to one claim and one source instead of an entire answer. If longer responses dominate, separate length from the quality rubric. If impressive source titles bias decisions, supply the relevant passage with its necessary context.

Unstable cases enter a human-review queue whose size is reported. Do not discard difficult examples and publish only easy-case accuracy. Record prompt revisions and attempts rather than silently selecting favorable outcomes. Model, prompt or threshold changes require rerunning a frozen calibration set.

Zhihe Growth's delivery boundary

Zhihe Growth can use model-assisted screening to flag potentially unsupported citations, omitted market restrictions or model-number confusion. Deliver evidence locations, reasons and unresolved states alongside scores. Increased brand presence is not improved technical correctness.

This material supplies procedures, negative-example design and worksheets. No new commercial-platform tests or independent human calibration were performed for it. A future execution must identify model, date, sampling configuration and adjudicated outcomes. The citation-metrics guide defines the measures; repeat testing handles answer variability.

Calibration exercise: audit the reason, not only the score

These are proposed controls, not completed commercial-model evaluations. Freeze evidence before testing each pair.

Control Fixed Changed Review focus
Order swap Both answers A/B position Reversal tied to position
Short versus long Facts and sources Redundancy No factual reward for repetition
Brand blinding Specifications Company name Evidence-based reasoning
Source removal Assertion Supporting material Missing support reflected
Unit corruption Other wording 20 L to 20 mL Exact conflict identified
Honest uncertainty Available evidence Explicit abstention No reward for guessing

An auditable reason identifies the claim, evidence and rule. “A sounds more professional” is insufficient. Permit both answers to fail rather than forcing a winner. A rationale identifying a critical factual error alongside a perfect score is itself a judge failure.

Download and primary references

Use the review and lab package to build a blinded set before connecting an evaluator. Do not transmit confidential customer material or unauthorized documents to an external review service.

Primary references are the judge study and scikit-learn's kappa documentation. This guide establishes a measurement discipline, not a guarantee that any model can independently accept factual work.

Knowledge center · GEO services · Research and evidence