Match the budget to the research question
Estimating current citation frequency, comparing site versions, finding critical factual errors and covering a new market are different tasks. Proportion estimation concerns precision; rare harmful errors need risk-focused sampling; comparison requires an effect-size and control design. One convenient question count cannot establish every objective.
Specify the population and question source, such as English unbranded procurement questions from European distributors. Separate branded prompts, support questions, countries and languages, with explicit weights. A reported aggregate should identify the population it represents rather than implying universal AI visibility.
Use approximation as planning, not proof
For independent binary observations, a familiar planning approximation is n = 1.96²p(1-p)/e², rounded upward. Here p is the assumed proportion and e the desired half-width.
| Assumed p | Half-width e | Approximate independent n |
|---|---|---|
| 0.1 | 0.05 | 139 |
| 0.1 | 0.10 | 35 |
| 0.5 | 0.05 | 385 |
| 0.5 | 0.10 | 97 |
These executed calculations predict neither this website's citation rate nor the power of a two-group comparison. Normal approximation can be inappropriate near zero or one, with small samples or complex designs. Report intervals suited to the design and explain whether assumptions came from historical data or a conservative planning choice.
Repetitions are not independent questions
Three responses to one question may share retrieval sources and platform preferences. A simple illustrative design effect is 1+(m-1)rho. With three repetitions and assumed within-question correlation 0.5, it equals two. This assumption was not estimated from a commercial platform.
More distinct questions broaden coverage; additional repeats reveal stability. Risky model-specific questions can receive extra checks, but repeating favorable questions must not inflate the aggregate. See the repeat-test protocol for question-level accounting.
Include missing data and review effort
Allow time for timeouts, inaccessible sources and account restrictions, but do not treat missing observations as freely disposable. Record attempts, verified answers, pending sources and failures. Disclose missingness with any verified-only denominator.
Source verification often costs more than sending a prompt: resolve URLs, read evidence, match assertions, check versions and adjudicate conflicts. Budget separately for prompting, source recovery, human review and exceptions. Call count alone is insufficient, and model self-scoring is not independent human acceptance.
Decide when to stop before seeing the score
A fixed plan can specify counts per stratum, repeats and an end date, followed by one summary. If cost, safety or platform limits cause an early stop, disclose the reason and completion fraction. Stopping as soon as the score looks favorable creates selective-reporting risk.
Sequential decisions require a prespecified method, inspection schedule and thresholds. The package does not implement sequential testing or endorse a universal citation target as a stopping rule. Its blank planning template records the primary metric, strata, fixed end date and missing-data policy.
Zhihe Growth's reporting responsibility
Zhihe Growth should deliver raw counts, scope and uncertainty rather than a percentage alone. Historical SuperPDR and ELEREIN figures retain their own definitions and are not guaranteed targets for a new client. For exporters, start with representative buyer questions and high-risk facts, then expand sampling to stabilize conclusions.
When outcomes disappoint, investigate evidence and access rather than changing denominators or deleting difficult questions. A claim about the effect of a redesign needs a separate attribution design. Planning improves interpretability, not guaranteed recommendations or revenue.
Planning exercise: complete the decision table first
| Field | Question | Acceptable record | Avoid |
|---|---|---|---|
| Primary metric | Citations, correctness or inquiries? | Defined beforehand | Select the biggest increase later |
| Sampling | How are questions included? | Demand and strata rules | Only previously cited questions |
| Repetitions | How many per question? | Fixed beforehand | Repeat until cited |
| Cost limit | When to stop and escalate? | Explicit budget | Unlimited retries |
| Missingness | How to handle unresolved URLs? | Pending and recovery policy | Delete unknowns |
| End condition | When to summarize? | Prespecified date or method | Stop on a favorable score |
When resources are limited, narrow the claim: English distributor questions do not represent all global buyers. Risk-focused oversampling is not the natural demand distribution. Use exploratory small samples to diagnose failures rather than rank platforms.
Later budgets may learn from error clusters and review cost, but preserve each new plan separately instead of rewriting the original success criterion.
Reproduction and references
The runnable package records the approximation, four outputs and assumed design effect in results.json. Read the interval guide and SciPy bootstrap documentation. This exercise performed local arithmetic, not new external AI tests.