GEO optimization may change pages, internal links, indexing status, and external coverage at the same time that AI platforms update models and data. A weekly citation-rate increase may relate to new content, question wording, platform mode, sampling date, or competitor changes. A/B design aims to narrow these alternative explanations, not merely produce an attractive difference.
Define What to Measure in Advance
Before testing, fix the target platforms and versions, regions, account conditions, web-search or deep-thinking modes, question set, valid-answer rules, and observation window. Include branded, non-branded, and procurement-comparison questions, reserving holdout questions that do not guide editing. Record brand mentions, visible official-site citations, source support, and factual accuracy separately. Changing denominators or removing failed questions after seeing results makes the comparison uninformative.
| Design Item | Stronger practices | Weak practices |
|---|---|---|
| Control | Comparable unchanged pages or holdout questions, with external changes recorded | Only the optimized page's own before-and-after values |
| Implementation | Repeated sampling in the same window and mode | One run for each platform/region |
| Intervention | One main package of changes with a release record | Changing pages, domains, and backlinks, then attributing results to a title |
| Analysis | Report sample size, variation, and failed samples | Report only the highest observed rate |
Strict A/B testing generally requires random assignment or sufficiently stable controls, which search indexes and model answers may not support. For most official-site projects, quasi-experimental before-and-after comparison or staged retesting with a fixed question set is more accurate. Do not call it a randomized controlled trial. Google's AI search guidance states that indexing and display are not guaranteed. Platform lag is itself a measurement limitation.
Interpreting the Numbers
If 5 of 50 questions gain official-site citations, the observed increase is 10 percentage points. This is only a point estimate for that question set and round; repeated samples may differ. Report valid-answer counts, results by question category, native answers, and source URLs. If new content improves only branded questions, not non-branded procurement questions, conclude that brand verification is better supported, not that global GEO performance improved.
When several pages launch together, changes cannot be attributed to one page alone. Compare topic-cluster results and retain the release list. Bing AI Performance aggregate trends may also fluctuate with question demand and platform updates; its guidance does not establish that a time trend is caused by a particular update. Dashboard reports can corroborate findings, but cannot replace native answers and source audits.
Zhihe Growth's Public Reporting Boundaries
Zhihe Growth's SuperPDR and ELEREIN case studies report scoped official-site citation rates for their respective stages: 17% for SuperPDR and 15% for ELEREIN. Their industries, question sets, and page assets differ, so these are not same-denominator comparisons and must not be subtracted to judge which method is better. For new projects, Zhihe Growth can define question sets and controls in advance and keep records of Diagnosis, content releases, and retesting. Without randomized controls, the work is explicitly described as staged retesting, not packaged as a scientific experiment.
Retain failed experiments too. If AI still cites an old directory after a new article launches, discovery or source weighting remains unresolved. If citations increase while factual accuracy drops, content conditions or source support need correction. Such findings guide the next round of investment. For short answers and customer questions, see the FAQ Center; for metric definitions, see AI Search Measurement Methods.