A citation rate needs a numerator, denominator and uncertainty statement. Results from different platforms, question sets and modes should not be casually pooled. The numerical examples here use synthetic inputs, not client retests.
Distinguish level, percentage points and relative change
A move from 10% to 15% is a five-percentage-point change, a 50% relative increase, and a final level of 15%. These describe different quantities. Without a baseline and comparable records, a reported 15% cannot establish a 50% increase.
A different range comes from unresolved sources. With K confirmed citations, U unresolved records and N valid answers, the bounds K/N to (K+U)/N describe incomplete classification. They are not a 95% confidence interval.
Wilson intervals for independent binary observations
Suppose x of n independent, consistently defined valid answers cite the official site, with p=x/n. Using z approximately 1.96, the Wilson interval adjusts the center and width to avoid some failures of a simple normal approximation for small samples or boundary proportions.
denom = 1 + z*z/n
center = (p + z*z/(2*n)) / denom
half = z * sqrt(p*(1-p)/n + z*z/(4*n*n)) / denom
interval = [center-half, center+half]
For x=2 and n=10, the interval is approximately 5.7% to 51.0%. The width signals limited precision. A zero denominator is not a zero rate: the program rejects it, along with negative successes or successes exceeding the sample size. See the program and package.
Repeated questions require a dependence model
Repeated answers to one question share intent, likely sources and platform conditions. Treating 50 questions asked three times as 150 independent questions can understate uncertainty. Platform changes and shared sessions may also create dependence across questions.
The teaching example treats each question as a cluster: calculate its mean after-minus-before difference, then resample whole paired question clusters. A real design may require additional date, platform or account structure. Question clustering is not a universal solution for every dependency.
Inspect the executable synthetic example
The data contains six questions with three repeats per period, moving from 3/18 to 9/18. The program uses seed 20261002 and 10,000 resamples. The observed output gives a paired difference of 0.3333 and a percentile interval from 0 to 0.6667: zero to about 66.7 percentage points.
This does not prove a significant improvement. Six clusters are too few for reliable real-world inference, the distribution is discrete, and external confounding is uncontrolled. The purpose is to demonstrate the resampling unit. See the SciPy bootstrap documentation for general methods; this package explicitly resamples question differences with NumPy.
Interval overlap is not a paired comparison test
Computing one interval before and one after, then checking whether they overlap, does not replace analysis of paired changes. The same questions create dependence. Independent samples require an appropriate independent comparison instead. Trying many subsets until a narrow interval appears introduces selection bias.
Specify the primary measure, groups and stopping rule before testing. Resolve pending sources, handle technical failures according to protocol and report platform strata. Confidence intervals do not correct bad labels or prove that a website change was the sole cause.
Separate unresolved-source bounds from sampling intervals
Test boundary behavior: x=0 of n=10 has a zero lower bound but a positive upper bound; x=n has an upper bound of one but a lower bound below one. Reject x greater than n and n=0. Use integer counts, not invented denominators reconstructed from rounded percentages, and round only for display.
Unequal repeat counts introduce a weighting choice. Averaging question-level rates weights questions equally; pooling runs weights frequently repeated questions more heavily. This example has three repeats per question, so the point estimates coincide, but that does not generalize. Report both distinct-question count and total runs.
Suppose ten valid answers include two confirmed official citations, three unresolved source links and five confirmed absences. The identified proportion ranges from 20% to 50%. Recovering sources can narrow this range; merely repeating more questions cannot repair the existing missing labels. Applying Wilson to two successes would incorrectly treat three unresolved answers as confirmed failures.
The earlier Wilson calculation assumes completed labels and independent binary observations. Six questions repeated three times need their question structure retained. Opening one source card three times is not three answers, and follow-ups in one conversation are not automatically independent sessions.
Do not use extra decimal places to disguise wide uncertainty. Report counts, point estimate, method and limits. Zero observed successes does not establish a permanently zero probability, just as all successes do not guarantee 100% future performance. Sample planning also requires a target difference and question population. More diverse fixed-intent questions and more repeats of one question provide different information. The six-cluster example cannot support a commercial performance guarantee.
How Zhihe Growth should report uncertainty
Zhihe Growth can show clients estimates, sample sizes, intervals, failures and limitations to support prioritization. It cannot attach invented intervals to SuperPDR's 17% or ELEREIN's 15% when the disclosed percentages alone do not establish original denominators and dependence structure.
Formal calculations require authorized original records, not guessed sample sizes. Continue with the metric dictionary, attribution guide and experiment design. Teaching calculations do not alter disclosed case definitions.