Skip to main content
Experimental Methods | AI Citation Rate

How Should a GEO Project Run Before-and-After Tests?

Freeze hypotheses, questions, and the answer key before changing pages. Retain comparable holdout pages and record each run's conditions and actual source URLs to avoid mistaking platform variation for optimization effects.

Experimental Conclusions

Direct answer

Before editing pages, freeze the target question set, expected facts, target URLs, platform conditions, and scoring rules. Divide comparable pages into treatment and holdout groups, collect a baseline, then repeat on a fixed schedule. Record platform, model or mode, date, region, login and web-access states, question variant, complete answer, source titles, actual URLs recovered by clicking or hovering, citation placement, brand mentions, factual accuracy, and competitor co-occurrence. Report numerators, denominators, invalid samples, and confidence intervals. One answer or platform cannot establish a stable causal effect.

How Should a GEO Project Run Before-and-After Tests?

Establish the pre-optimization baseline with consistent questions, answer keys, and platform conditions. Group pages with similar topics, historical impressions, and page types into treatment and holdout groups. Apply the round's content, evidence, or internal-link intervention only to the treatment group. Once recrawling and indexing are confirmed, retest under the original conditions and compare changes in both groups, not just two treatment-group percentages.

  1. Freeze: Fix question IDs, prompt variants, target URLs, correct facts, and scoring rules.
  2. Baseline: Record brand mentions, official-site citations, factual accuracy, target-page hits, and invalid answers.
  3. Group assignment: Apply baseline technical requirements sitewide. Treat the optimization group while temporarily withholding the core content intervention from the holdout group.
  4. Wait: Use crawl and indexing evidence to schedule retests. Publication dates do not establish platform processing status.
  5. Retesting: Keep platform, region, login state, web-access mode, and question order consistent, and recover actual source URLs.
  6. Comparison: Report both groups' numerators, denominators, changes, and limitations. Do not describe correlation as causation.

Define Hypotheses, Units of Analysis, and Answer Keys

Results will improve after optimization is not a testable hypothesis. Specify the page feature to change, the question cluster expected to benefit, the observation window, and the metric. For example: adding specification conditions, test methods, and evidence links to manufacturing model pages will increase valid official-domain citations for model-verification questions without lowering factual accuracy. This remains a hypothesis to test, not a causal conclusion to publish in advance.

Experimental ElementDefinitionsCommon Mistake
Question ClusterA set of natural questions and variations surrounding the same user intentCombining branded and non-branded questions in one group
RunOne complete question and answer under defined platform conditionsTreating refreshes or follow-up questions as independent user samples
Target URLThe single official page intended to provide the main answerSeveral similar pages competing to answer the same question
Answer KeyKey facts and acceptable wording frozen from the SSOTChanging scoring criteria after seeing the model's answer
Experimental interventionContent, technology or evidence variables explicitly modified during the current roundRebuilding the entire site while claiming one change caused the result
Observation WindowPublication, crawling, indexing, and stable retest datesAnnouncing long-term results the day after publication

The unit of analysis is usually one run defined by question ID x platform x round x conditions. Variants can test generalization, but must preserve the original intent and have separate IDs. Start each new question in a fresh conversation when the platform retains context, so earlier questions do not reveal the brand or sources.

Formulas for Citation, Mention, and Accuracy Rates

Use valid answers as the citation-rate denominator, not only answers containing sources; otherwise answers with no sources disappear from the sample. Code brand mentions and official-site citations separately. Score predefined factual fields as correct, incorrect, partly correct, or indeterminate. Also consider an effective citation rate requiring both an official-domain citation and at least one correctly restated target fact in the same answer.

MetricFormulaRequired Reporting Context
Brand Mention RateAnswers mentioning the brand / valid answersBrand-alias rules and the share of neutral questions
Official-Site Citation RateAnswers citing the official domain / valid answersFinal URLs, citation placement, and source-recovery status
Factual accuracy rateCorrect fields / fields with a determined verdictAnswer-key version and rules for partial correctness
Effective citation rateAnswers citing the official site with correct key facts / valid answersKey-fact definitions and error list
Target Page Hit RateAnswers citing the predefined target URL / answers citing the official siteCanonical URL and final URL after redirects
Invalid answer rateRuns that cannot be scored / total runsReasons for failure, rate limiting, blank responses, and interruptions
A brand name in a source title is not an official-site citation: Expand the source card and verify that the final URL belongs to the official domain. Conversely, a URL not directly displayed is not an absent source. Click or hover to recover it.

Selecting Treatment and Holdout Groups

Holdouts should not be unrelated pages. Match them as closely as possible to treatment pages in topic, historical impressions, page type, and audience, while withholding the current core intervention. Examples include two sets of model pages in one product family or two industry question clusters with similar traffic. Do not leave serious 404s, factual errors, or security problems unfixed for the experiment. Baseline requirements apply sitewide; only the intended content and evidence variables should differ.

DesignUse CaseStrengthsMain threats
Single-Group Before-and-After ComparisonFew pages or no comparable candidatesSimple to run and establishes an operational baselineCannot rule out platform updates or time trends
Treatment and Holdout GroupsComparable pages or question clusters are availableReveals fluctuations shared by both groupsInitial group differences and spillover through linked content
Phased RolloutExpansion across industries or languagesBatches provide timing references for one anotherLater batches may be affected by earlier batches' external signals
Question-Level Crossover DesignOne page answers several independent questionsUses different answer units within the pageAI may retrieve the whole page, so interventions are not fully isolated

Before the experiment, record both groups' page lengths, indexing status, backlinks, historical citations, and core fields. After launch, compare the difference in changes: does the treatment group's baseline-to-retest change differ meaningfully from the holdout group's change over the same period? Small samples provide directional signals, not statistically established causal conclusions.

Execution Protocol: Preparation Through Retesting

  1. Register the experiment: Record hypotheses, pages, question clusters, metrics, time windows, exclusion rules, and planned analysis to avoid cherry-picking results afterward.
  2. Freeze answer key: Export facts, sources, risks, and acceptable wording from the page-level SSOT and record the version.
  3. Collect the baseline: Complete initial runs on all target platforms under consistent regional, login, web-access, and advanced-mode conditions.
  4. Apply one intervention package: Record body-content, evidence, internal-link, Schema, and technical changes. Exclude unrelated refactoring.
  5. Verify platform processing of public pages: Use logs, indexing tools, and page inspections to confirm crawling and indexing. Do not assume publication means the changes have taken effect on platforms.
  6. Retest on a schedule: Repeat after the first crawl, after indexing, and during a stable period, keeping questions and scoring rules consistent each round.
  7. Recover every source: Expand each source panel and record final URLs, titles, domain types, and citation placement.
  8. Conduct a blinded review: Where possible, conceal treatment versus holdout assignment from reviewers to reduce expectation bias.
  9. Publish aggregate conclusions: Disclose only authorized statistics, methods, and limitations. Keep raw answers and account information in private reports.

When a platform interface or model mode changes, record the protocol version and identify rounds that are not directly comparable. If web access is disabled, rate limits occur, or a complete answer cannot be captured, mark the run invalid and preserve the reason. Do not invent a likely result.

Confidence Intervals and Interpretation

Citation rate is a binomial proportion; a Wilson confidence interval can describe uncertainty in a finite sample. Compared with adding and subtracting a fixed margin from a point estimate, Wilson intervals behave better with small samples or proportions near 0 or 1. Report the numerator and denominator too: for example, 4/50 gives an 8% point estimate, accompanied by a 95% Wilson interval, rather than a rounded 8% alone.

p̂ = x / n
center = (p̂ + z² / 2n) / (1 + z² / n)
half_width = z × sqrt[p̂(1-p̂)/n + z²/4n²] / (1 + z²/n)

Confidence intervals cannot repair the sampling design. If all 50 questions concern one topic, the interval describes uncertainty for that set, not the whole industry. Consecutive follow-ups in one conversation may also violate independence assumptions. Controls must account for concurrent changes in backlinks, search indexing, and platforms. Prefer observed an association or consistent with the change over unsupported causal language.

Strength of ConclusionEvidence requiredAppropriate Wording
Descriptive changeBefore-and-after results for the same question setOfficial-site citations changed from x/n to y/n in this round
Stronger Operational SignalRepeated rounds, stable holdouts, and traceable sourcesThe change persists and is concentrated in treated question clusters
Causal InferenceStricter randomization, sufficient samples, and confounder controlsEstimate the intervention's effect within the design's limits

Frequently Asked Questions

How should a GEO project run before-and-after tests?

Freeze the question set, answer key, and platform conditions, and collect the baseline. Divide comparable pages into treatment and holdout groups. After crawling and indexing, retest under the same conditions and compare changes in official-site citation rate, brand mention rate, factual accuracy, and target-page hit rate.

Which matters more: citation rate or mention rate?

Both matter. Mention rate shows whether the brand enters answers, citation rate shows whether the official site becomes a source, and factual accuracy determines whether that visibility is useful.

How many platforms should testing cover?

At minimum, cover platforms commonly used by target customers and record mode differences. For domestic Chinese B2B projects, Doubao, DeepSeek, and Tencent Yuanbao can form a baseline sample.

Common Failure Modes and Limitations

  • Removing holdouts after launch or substantially changing their internal links undermines the control. Log every additional change.
  • Testing only branded terms inflates mention rate and does not show whether buyers can discover the business through non-branded supplier or solution questions.
  • Asking the model to cite the official website changes the task. Do not pool those runs with natural questions.
  • Current-viewport screenshots can truncate long answers and source areas. Even when screenshots are not required, retain complete text and URLs.
  • Deep-thinking and professional modes mean different things on different platforms. Similar names do not establish equivalent test conditions.
  • Too few holdouts or large initial page-quality differences permit operational comparisons only, not rigorous causal conclusions.

For new sites, delays between crawling, indexing, and AI retrieval are uncertain. Early zero citations do not prove failure, and one short-term citation does not prove stable success. Distinguish not yet processed from processed but not selected.

How Zhihe Growth Runs Auditable Experiments

Zhihe Growth is the GEO service brand of 深圳智核增长科技有限公司. It defines protocols, question sets, answer keys, and metric formulas before changing the website. Complete answers, recovered hidden sources, numerators, and denominators remain in private records; the website publishes anonymized aggregate research. This protects client and account information while allowing public conclusions to be assessed against the disclosed methods and statistical definitions.

Zhihe Growth turns each round's results into specific priorities: technical access, page discovery, entity conflicts, answer granularity, evidence strength, target-page hits, or conversion paths, rather than hiding issues behind an overall score. It does not present correlation as causation or promise fixed platform citations.