Skip to main content
Measurement Tools | Microsoft Bing

Bing AI Performance Public Preview: How Businesses Can Read AI Citation Data

Bing Webmaster Tools began giving site owners visibility into some citation activity in AI answers. The dashboard adds an official data source for GEO measurement, but businesses must understand its sampling, aggregation and metric definitions before using it to assess project delivery.

Key Takeaway

Bing AI Performance helps businesses observe total citations, average daily cited pages, grounding queries and URL-level citation trends across selected Microsoft AI experiences. It is not an AI answer ranking table and does not establish page authority, placement or conversions. Combine it with fixed-question tests, page-change records, organic search data and server logs to build a measurement system supported by multiple sources of evidence.

What's the official change?

On February 10, 2026, Microsoft Bing announced the public preview of AI Performance in Bing Webmaster Tools. It described an aggregate view of site citations in AI answers as a step toward greater transparency in generative search and the open web. The announced coverage included Microsoft Copilot, AI-generated summaries in Bing and selected partner integrations; coverage could change as the product developed.

The dashboard provides four main measures. Total Citations counts citations displayed as sources in AI answers during the selected period. Average Cited Pages shows the average daily number of unique site pages displayed as sources. Grounding queries provides a sample of key phrases used in AI retrieval and associated with cited content. Page-level citation activity aggregates citations by URL. Trend charts help site owners track changes over time.

Bing specifies limits for each metric: citation counts do not indicate answer placement; average cited pages does not measure ranking or authority; page-level activity does not establish page importance. Grounding queries is a sample, not a complete query log. Rising numbers therefore cannot simply be called improved AI rankings, nor can absolute values be compared across sites without accounting for the time window, site scale and coverage.

Publication date 2026-02-10
What it measures Source citation activity in AI answers
What it does not establish Ranking, authority or answer position

Implications for Brands Expanding Internationally

For businesses with English-language sites, international product sites or multilingual knowledge centers, the dashboard provides data previously gathered mainly through screenshots: which pages become AI sources, which retrieval topics surface site content and whether citation trends persist after updates. Product specifications, integration documentation, troubleshooting guides, industry comparisons and evidence pages were difficult to assess through traditional click reports alone. URL-level citation activity provides an additional, partial view of their use in AI answers.

It also changes content evaluation. Look beyond which article receives the most citations to whether cited pages cover the decision process: brand discovery, capability checks, limitations, evidence, service comparisons and contact actions. If only definition pages are cited while core product or service pages remain absent, knowledge content may lack explanatory links to the business. Citations concentrated on outdated URLs call for checks of canonicals, redirects, the update log and IndexNow notifications.

For multilingual sites, URL-level data can help identify whether similar grounding queries surface pages for different markets. If Chinese pages receive citations but English pages show no activity, do not immediately blame translation quality. Check English URL indexing, robots and WAF consistency, hreflang, the independent value of the page's facts and the platform's regional coverage.

Zhihe Growth's Assessment

Bing AI Performance is valuable because it makes site citations partially verifiable through official data, not because it adds another score. Use it for trend evidence and page discovery, not as the sole basis for accepting project delivery. Zhihe Growth recommends four measurement layers: official dashboards for aggregate trends; fixed question sets for repeatable brand-mention and source checks; server logs for AI bot crawling; and business analytics for qualified visits and inquiries.

Use grounding queries to maintain question sets, but do not create one page for every query mechanically. Group related phrases by intent, then assess whether existing articles, FAQs or service pages answer them adequately. For relevant pages with few citations, improve direct answers, titles, field tables, evidence and update dates. Do not expand into unrelated topics merely to chase dashboard phrases.

Bing recommends IndexNow notifications for new, updated or deleted URLs. Zhihe Growth treats this as a technical publishing step, not an indexing guarantee. Notifications are useful only when pages are accessible, have correct canonicals, offer valuable content and are not blocked by robots or WAF rules. Send update notifications for substantive changes, not repeated submissions of unchanged pages.

How to Interpret the Four Metrics

MetricWhat it can tell youWhat it cannot tell you
Total CitationsWhether total recorded activity showing the site as an AI answer source changed during the selected period.Where citations appear in answers, whether users click them or whether they outperform competitors.
Average Cited PagesWhether the daily number of unique cited pages on the site expanded or contracted.The authority, ranking or role of each page in the answer.
Grounding queriesWhich sampled retrieval phrases relate to site citations and can inform question clusters.All user questions, a complete retrieval log or a definitive distribution of user intent.
Page-level activityWhich URLs have been cited more often, and whether trends have changed before and after the update.Page importance, answer placement, conversion quality or causality.

GEO Checklist for International Brands

  1. Freeze the pre-optimization baseline. Export total citations, average daily cited pages, query samples and URL activity for a consistent time window. Record the account, region and export date.
  2. Label page types. Classify cited URLs as service, product, technical article, FAQ, evidence, case study or company entity pages to assess coverage of the full decision process.
  3. Create a page holdout group. Select pages with similar topics and historical performance. Optimize structure and evidence for one group and leave the other unchanged, so site-wide fluctuations are not mistaken for single-page effects.
  4. Record actual publishing changes. Keep records of body text, titles, Schema, canonicals, internal links, dates and IndexNow notification times so trends can be investigated.
  5. Run supplementary manual question tests. Create natural-language variants of grounding queries. Record the platform, date, login state, brand mentions, citations, source URLs and factual accuracy.
  6. Connect search and operational data. Also examine Bing organic impressions and clicks, Chat/Copilot referrals, landing-page engagement and inquiries. Citation counts are not customer acquisition counts.
  7. Review monthly instead of drawing daily conclusions. Public preview data may have processing delays and sample fluctuations, and sufficient time windows should be used to observe direction.

Limitations

At the time of the February 2026 announcement, AI Performance was in public preview. Supported AI experiences, metric definitions, sampling and processing could change. The cited help documentation described daily refreshes with processing delays and aggregate or sampled data, not complete audit logs of every answer. The dashboard did not provide full context for citation text, answer placement, clicks or competing sources shown alongside it.

Citation growth establishes only that recorded citation activity increased. Alone, it proves neither that content changes caused the increase nor that trust, organic rankings or sales improved. Formal project evaluation requires documented time windows, control pages and repeated tests across platforms.

Official sources