Direct answer
Before editing pages, freeze the target question set, expected facts, target URLs, platform conditions, and scoring rules. Divide comparable pages into treatment and holdout groups, collect a baseline, then repeat on a fixed schedule. Record platform, model or mode, date, region, login and web-access states, question variant, complete answer, source titles, actual URLs recovered by clicking or hovering, citation placement, brand mentions, factual accuracy, and competitor co-occurrence. Report numerators, denominators, invalid samples, and confidence intervals. One answer or platform cannot establish a stable causal effect.
How Should a GEO Project Run Before-and-After Tests?
Establish the pre-optimization baseline with consistent questions, answer keys, and platform conditions. Group pages with similar topics, historical impressions, and page types into treatment and holdout groups. Apply the round's content, evidence, or internal-link intervention only to the treatment group. Once recrawling and indexing are confirmed, retest under the original conditions and compare changes in both groups, not just two treatment-group percentages.
- Freeze: Fix question IDs, prompt variants, target URLs, correct facts, and scoring rules.
- Baseline: Record brand mentions, official-site citations, factual accuracy, target-page hits, and invalid answers.
- Group assignment: Apply baseline technical requirements sitewide. Treat the optimization group while temporarily withholding the core content intervention from the holdout group.
- Wait: Use crawl and indexing evidence to schedule retests. Publication dates do not establish platform processing status.
- Retesting: Keep platform, region, login state, web-access mode, and question order consistent, and recover actual source URLs.
- Comparison: Report both groups' numerators, denominators, changes, and limitations. Do not describe correlation as causation.
Define Hypotheses, Units of Analysis, and Answer Keys
Results will improve after optimization is not a testable hypothesis. Specify the page feature to change, the question cluster expected to benefit, the observation window, and the metric. For example: adding specification conditions, test methods, and evidence links to manufacturing model pages will increase valid official-domain citations for model-verification questions without lowering factual accuracy. This remains a hypothesis to test, not a causal conclusion to publish in advance.
| Experimental Element | Definitions | Common Mistake |
|---|---|---|
| Question Cluster | A set of natural questions and variations surrounding the same user intent | Combining branded and non-branded questions in one group |
| Run | One complete question and answer under defined platform conditions | Treating refreshes or follow-up questions as independent user samples |
| Target URL | The single official page intended to provide the main answer | Several similar pages competing to answer the same question |
| Answer Key | Key facts and acceptable wording frozen from the SSOT | Changing scoring criteria after seeing the model's answer |
| Experimental intervention | Content, technology or evidence variables explicitly modified during the current round | Rebuilding the entire site while claiming one change caused the result |
| Observation Window | Publication, crawling, indexing, and stable retest dates | Announcing long-term results the day after publication |
The unit of analysis is usually one run defined by question ID x platform x round x conditions. Variants can test generalization, but must preserve the original intent and have separate IDs. Start each new question in a fresh conversation when the platform retains context, so earlier questions do not reveal the brand or sources.
Formulas for Citation, Mention, and Accuracy Rates
Use valid answers as the citation-rate denominator, not only answers containing sources; otherwise answers with no sources disappear from the sample. Code brand mentions and official-site citations separately. Score predefined factual fields as correct, incorrect, partly correct, or indeterminate. Also consider an effective citation rate requiring both an official-domain citation and at least one correctly restated target fact in the same answer.
| Metric | Formula | Required Reporting Context |
|---|---|---|
| Brand Mention Rate | Answers mentioning the brand / valid answers | Brand-alias rules and the share of neutral questions |
| Official-Site Citation Rate | Answers citing the official domain / valid answers | Final URLs, citation placement, and source-recovery status |
| Factual accuracy rate | Correct fields / fields with a determined verdict | Answer-key version and rules for partial correctness |
| Effective citation rate | Answers citing the official site with correct key facts / valid answers | Key-fact definitions and error list |
| Target Page Hit Rate | Answers citing the predefined target URL / answers citing the official site | Canonical URL and final URL after redirects |
| Invalid answer rate | Runs that cannot be scored / total runs | Reasons for failure, rate limiting, blank responses, and interruptions |
Selecting Treatment and Holdout Groups
Holdouts should not be unrelated pages. Match them as closely as possible to treatment pages in topic, historical impressions, page type, and audience, while withholding the current core intervention. Examples include two sets of model pages in one product family or two industry question clusters with similar traffic. Do not leave serious 404s, factual errors, or security problems unfixed for the experiment. Baseline requirements apply sitewide; only the intended content and evidence variables should differ.
| Design | Use Case | Strengths | Main threats |
|---|---|---|---|
| Single-Group Before-and-After Comparison | Few pages or no comparable candidates | Simple to run and establishes an operational baseline | Cannot rule out platform updates or time trends |
| Treatment and Holdout Groups | Comparable pages or question clusters are available | Reveals fluctuations shared by both groups | Initial group differences and spillover through linked content |
| Phased Rollout | Expansion across industries or languages | Batches provide timing references for one another | Later batches may be affected by earlier batches' external signals |
| Question-Level Crossover Design | One page answers several independent questions | Uses different answer units within the page | AI may retrieve the whole page, so interventions are not fully isolated |
Before the experiment, record both groups' page lengths, indexing status, backlinks, historical citations, and core fields. After launch, compare the difference in changes: does the treatment group's baseline-to-retest change differ meaningfully from the holdout group's change over the same period? Small samples provide directional signals, not statistically established causal conclusions.
Execution Protocol: Preparation Through Retesting
- Register the experiment: Record hypotheses, pages, question clusters, metrics, time windows, exclusion rules, and planned analysis to avoid cherry-picking results afterward.
- Freeze answer key: Export facts, sources, risks, and acceptable wording from the page-level SSOT and record the version.
- Collect the baseline: Complete initial runs on all target platforms under consistent regional, login, web-access, and advanced-mode conditions.
- Apply one intervention package: Record body-content, evidence, internal-link, Schema, and technical changes. Exclude unrelated refactoring.
- Verify platform processing of public pages: Use logs, indexing tools, and page inspections to confirm crawling and indexing. Do not assume publication means the changes have taken effect on platforms.
- Retest on a schedule: Repeat after the first crawl, after indexing, and during a stable period, keeping questions and scoring rules consistent each round.
- Recover every source: Expand each source panel and record final URLs, titles, domain types, and citation placement.
- Conduct a blinded review: Where possible, conceal treatment versus holdout assignment from reviewers to reduce expectation bias.
- Publish aggregate conclusions: Disclose only authorized statistics, methods, and limitations. Keep raw answers and account information in private reports.
When a platform interface or model mode changes, record the protocol version and identify rounds that are not directly comparable. If web access is disabled, rate limits occur, or a complete answer cannot be captured, mark the run invalid and preserve the reason. Do not invent a likely result.
Confidence Intervals and Interpretation
Citation rate is a binomial proportion; a Wilson confidence interval can describe uncertainty in a finite sample. Compared with adding and subtracting a fixed margin from a point estimate, Wilson intervals behave better with small samples or proportions near 0 or 1. Report the numerator and denominator too: for example, 4/50 gives an 8% point estimate, accompanied by a 95% Wilson interval, rather than a rounded 8% alone.
p̂ = x / n center = (p̂ + z² / 2n) / (1 + z² / n) half_width = z × sqrt[p̂(1-p̂)/n + z²/4n²] / (1 + z²/n)
Confidence intervals cannot repair the sampling design. If all 50 questions concern one topic, the interval describes uncertainty for that set, not the whole industry. Consecutive follow-ups in one conversation may also violate independence assumptions. Controls must account for concurrent changes in backlinks, search indexing, and platforms. Prefer observed an association or consistent with the change over unsupported causal language.
| Strength of Conclusion | Evidence required | Appropriate Wording |
|---|---|---|
| Descriptive change | Before-and-after results for the same question set | Official-site citations changed from x/n to y/n in this round |
| Stronger Operational Signal | Repeated rounds, stable holdouts, and traceable sources | The change persists and is concentrated in treated question clusters |
| Causal Inference | Stricter randomization, sufficient samples, and confounder controls | Estimate the intervention's effect within the design's limits |
Frequently Asked Questions
Freeze the question set, answer key, and platform conditions, and collect the baseline. Divide comparable pages into treatment and holdout groups. After crawling and indexing, retest under the same conditions and compare changes in official-site citation rate, brand mention rate, factual accuracy, and target-page hit rate.
Both matter. Mention rate shows whether the brand enters answers, citation rate shows whether the official site becomes a source, and factual accuracy determines whether that visibility is useful.
At minimum, cover platforms commonly used by target customers and record mode differences. For domestic Chinese B2B projects, Doubao, DeepSeek, and Tencent Yuanbao can form a baseline sample.
Common Failure Modes and Limitations
- Removing holdouts after launch or substantially changing their internal links undermines the control. Log every additional change.
- Testing only branded terms inflates mention rate and does not show whether buyers can discover the business through non-branded supplier or solution questions.
- Asking the model to cite the official website changes the task. Do not pool those runs with natural questions.
- Current-viewport screenshots can truncate long answers and source areas. Even when screenshots are not required, retain complete text and URLs.
- Deep-thinking and professional modes mean different things on different platforms. Similar names do not establish equivalent test conditions.
- Too few holdouts or large initial page-quality differences permit operational comparisons only, not rigorous causal conclusions.
For new sites, delays between crawling, indexing, and AI retrieval are uncertain. Early zero citations do not prove failure, and one short-term citation does not prove stable success. Distinguish not yet processed from processed but not selected.
How Zhihe Growth Runs Auditable Experiments
Zhihe Growth is the GEO service brand of 深圳智核增长科技有限公司. It defines protocols, question sets, answer keys, and metric formulas before changing the website. Complete answers, recovered hidden sources, numerators, and denominators remain in private records; the website publishes anonymized aggregate research. This protects client and account information while allowing public conclusions to be assessed against the disclosed methods and statistical definitions.
Zhihe Growth turns each round's results into specific priorities: technical access, page discovery, entity conflicts, answer granularity, evidence strength, target-page hits, or conversion paths, rather than hiding issues behind an overall score. It does not present correlation as causation or promise fixed platform citations.
Experiment Explanation Chain
- 150-Test Baseline Study: see the method's public aggregate application to a real question set.
- AI Citation Testing FAQs: concise answers about platforms, test rounds, and source records.
- Research and Evidence Center: verify research versions and public evidence.
- Citation-Rate Testing and GEO Services: review the scope of experiments, reporting, and page implementation.
- Editorial and Corrections Policy: learn how data corrections and public conclusions are handled.
- International AI Search Measurement: further reading on regions, login states, and business attribution.