Skip to main content
SEO 14 min

Bing AI Performance: Measure Copilot Citations | Ighenatt

A practical framework for measuring Copilot citations with Bing AI Performance, connecting citation signals to search data, referral traffic, and useful acti...

EG

Elu Gonzalez

Author

How do you measure Copilot citations with Bing AI Performance?

Use Bing Webmaster Tools AI Performance to track total citations, cited pages, grounding queries, and citation trends across its supported AI experiences. Treat a citation as evidence that a page was used as a source, then compare it with Bing Search Performance, referral sessions, and conversions. A citation is not a ranking, a guaranteed click, or proof of business impact.

Key takeaways

  • Bing Webmaster Tools AI Performance reports citations, cited pages, grounding queries, and trends across supported AI experiences; it is not a classic ranking report.
  • Total citations and average cited pages describe selection as a source, not the placement, prominence, or authority of a URL in an individual answer.
  • Grounding queries are a sample of retrieval phrases, not a complete record of every user prompt, so use them to form hypotheses rather than to claim full query coverage.
  • Pair AI Performance with Bing Search Performance and analytics: citations show source selection, while clicks, sessions, and conversions describe different parts of the journey.
  • Prioritise pages that are already cited or repeatedly retrieved, verify factual claims, and monitor a fixed period before deciding whether an edit changed performance.

Measure Copilot in terms of the pages cited, the retrieval phrases associated with them, their frequency, and what happens after a person follows a source link.

That question now has a first-party starting point. Microsoft introduced AI Performance in Bing Webmaster Tools as a public preview on 10 February 2026. Microsoft’s announcement says it reports citation activity for publisher content across Microsoft Copilot, AI-generated summaries in Bing, and selected partner integrations. The product is still a preview, so treat its current coverage and definitions as product documentation, not as a permanent reporting contract.

This is a different job from tracking a blue-link position. A Copilot citation tells you that a URL was displayed as a source in a generated answer during a selected period. It does not tell you that the URL was first, persuasive, clicked, or commercially valuable. Microsoft explicitly says that its citation counts do not show placement or presentation in an individual answer, and that page-level citation activity is not a measure of page importance, ranking, or authority. Microsoft’s metric definition gives the same boundary. Keep it visible in every AI visibility report.

This guide shows how to use the available data without turning it into a made-up funnel. It covers the reporting vocabulary, a repeatable measurement process, what to compare with Bing Search Performance and analytics, and the decisions that the data can actually support.

Start with the right unit of measurement

A Copilot citation is an observed source selection, not a rank. In the current AI Performance documentation, total citations are the number of times sources from a site were displayed in AI-generated answers within the selected period. Average cited pages is the average number of unique site pages shown as sources per day. The dashboard also provides grounding queries, page-level citation activity, and a citation trend. Microsoft documents each definition and its limits.

That gives the report a clean hierarchy:

Signal What it describes What it does not prove
Total citations How often your content was displayed as a source in supported AI answers during the period Rank, visibility position, click volume, or revenue
Average cited pages The daily average number of unique pages from your domain used as sources That a larger number means higher authority or better conversion
Grounding queries Key retrieval phrases used when content that cited your site was found The exact prompt a user typed, or a complete query log
Cited pages The URLs from your site that received citations That the cited sentence was the main answer or a recommendation
Bing Search Performance Bing search clicks, impressions, CTR, positions, and other documented source data A direct explanation for every AI citation
Referral sessions and conversions Visits and on-site behaviour captured by your analytics implementation The full influence of an answer that led to a later branded search

There is a reason to keep the terms this strict. Microsoft says the grounding-query data represents a sample of citation activity, and it says average cited pages are aggregated across supported AI surfaces. If a team writes “Copilot users searched for X” when the dashboard actually shows a sampled grounding phrase, it has swapped a retrieval signal for a user-behaviour claim. The numbers may look more decisive, but the conclusion has become weaker.

The interface itself makes citations worth tracking. Microsoft says Copilot Search in Bing presents cited sources and offers links to sources used for the generated answer; Bing’s general search documentation also says generative answers based on search results include references to source websites so users can verify them. Copilot Search in Bing and Microsoft’s explanation of Bing results describe the user-facing citation model. A source link can create a visit. It can also create awareness without a visit. The dashboard only measures the first part: that a source was shown.

What AI Performance can answer today

Open Bing Webmaster Tools, choose the verified site, then review the AI Performance area for the same date range you will use elsewhere. Before exporting anything, write the date range, country or device filters if the interface provides them, and the day you retrieved the data. A screenshot without this context is not a useful baseline.

The dashboard answers five operational questions.

1. Is the site being cited at all?

Start with total citations and the trend. This is a discovery check, not a scorecard. If the count is zero, check the date range, the property, and whether the relevant pages are indexable before deciding that the content has no AI presence. If it is non-zero, record the number and move to the pages rather than celebrating the site-level total.

Microsoft defines total citations as sources displayed in AI-generated answers during the selected timeframe. It does not promise that every citation opportunity across every Copilot experience is included, and the product is labelled public preview. The defensible wording is “Bing Webmaster Tools recorded N citations in its supported AI experiences for this period,” not “our brand appeared N times in Copilot everywhere.” Read the current scope and metric definition.

2. Which URLs are the source candidates?

Use the cited-pages view to identify the URLs that keep appearing. Page-level citation activity is where a site-wide metric becomes a content decision. Look for three patterns:

  • A small group of pages receives most citations. Audit these pages first, because they reveal what the system already finds usable.
  • A page is cited for a topic that is broader or narrower than its intended purpose. Check the heading structure, definition, evidence, and internal links before expanding it.
  • A historically important page is absent while a weak or outdated page is cited. Check canonicals, redirects, indexability, duplicated intent, and whether the current page actually answers the retrieval need.

Do not infer a winner from a single daily spike. Citation counts are observations across a selected timeframe, and an individual answer can use several sources. A URL with ten citations may have been a supporting reference in ten answers; it has not necessarily been the preferred recommendation ten times. Microsoft repeats that a citation count is not rank or prominence. The same limitation is stated in the Clarity Citation dashboard documentation, which also treats page citations as references rather than traditional rankings.

3. What retrieval language is associated with citations?

Grounding queries are useful because they expose some of the language used to retrieve your pages. They are not a replacement for keyword research and they are not the exact prompts that a person typed. Microsoft says they are key phrases the AI used to retrieve content and that the reported data is only a sample of overall citation activity. That limitation appears in the AI Performance metric definitions.

Use them as a diagnostic bridge:

  1. Export or record the grounding phrase and the cited URL.
  2. Search the URL for the phrase, related entities, and the answer promised by the page title.
  3. Read the passage a visitor would land on. Does it answer a narrow question, give a dated fact, explain a process, or merely mention the topic?
  4. Compare it with the page’s Bing search queries and with the questions your sales or support team receives.
  5. Make one hypothesis. For example: “This implementation guide is cited for migration checklists, but its checklist is three years old and no longer includes the current deployment step.”

One phrase may be enough to find a mismatch. It is not enough to claim that the phrase has a known search volume, that it is the original user query, or that changing the page will raise citations. Keep the hypothesis tied to the evidence you have.

4. Is the cited-page set widening or narrowing?

Average cited pages complements the total. Imagine that citations rise while the number of cited URLs falls: a small number of pages may be repeatedly selected. If citations fall while average cited pages rise, the site may be receiving scattered, low-frequency references. Those are descriptions, not diagnoses. You still need to inspect the specific pages and retrieval phrases.

This comparison is useful because it prevents a common reporting error: treating every increase in citation volume as broader topical coverage. Microsoft defines average cited pages as a daily average of unique URLs and warns that it does not indicate page authority or the page’s role in an answer. Use the product’s metric definition before attaching a business meaning.

5. Does the trend justify investigation?

The trend chart helps you choose where to look; it does not tell you why the line moved. Add annotations for page releases, major rewrites, indexing changes, migrations, outages, seasonal campaigns, and measurement changes. Then compare equivalent periods. A seven-day rise after a page edit is interesting. It is not causal proof, particularly when the report is a preview and AI answers vary by query and context.

For a modest site, a monthly review is usually more useful than chasing daily fluctuations. Keep a short change log next to the export. The discipline is simple: if you cannot say what changed on the site or in the measurement setup, do not write an explanation for the chart.

Build a measurement routine that survives scrutiny

A usable Copilot-citation report joins four distinct layers: AI citations, Bing search activity, on-site referrals, and business outcomes. Each layer answers a different question. Combining them without labels produces impressive-looking dashboards and poor decisions.

Step 1: Fix the baseline

Choose a stable period, such as the previous 28 complete days. Export or record total citations, average cited pages, cited URLs, grounding queries, and the date range. Save the raw export in the team’s analytics workspace, not only in a slide deck. Note that the AI Performance report is in public preview and supported surfaces may evolve.

At the URL level, create a small table with the cited page, citation count, associated grounding queries, last substantive update date, content owner, and the next review decision. Do not put a speculative “AI rank” column in it. The product does not supply that metric.

Step 2: Compare with Bing Search Performance, but do not merge the events

Bing Webmaster Tools’ standard Search Performance documentation defines impressions and clicks across several Bing locations. It specifically lists Chat among the sources and says that Web and Chat can report clicks, impressions, CTR, and average position; its keyword and page tables, however, are described as Web traffic rather than Chat data. Bing’s Search Performance help page sets out those source definitions.

This creates a practical comparison:

  • For each cited URL, inspect Bing search clicks, impressions, CTR, and documented query data over the same period.
  • Flag pages with citations but weak organic Bing discovery. They may answer a specific AI retrieval need that the current web-query report does not expose well.
  • Flag pages with strong Bing impressions but no citations. Check whether they are too broad, stale, duplicate, hard to parse, or simply not relevant to the sampled AI responses.
  • Keep “citation” and “click” in separate columns. A citation is displayed in an AI answer; a click is an interaction with a link. They are not interchangeable.

The comparison is directional, not a join key. The reports can cover different sources, metrics, and aggregation rules. Bing’s own documentation notes that Search Performance source reporting includes multiple places across Bing, while AI Performance aggregates citation activity across supported AI experiences. The data can sit beside each other, but it should not be summed into a fictional total “AI reach.”

Step 3: Measure the visit after the citation separately

Analytics can tell you what happens after a measurable visit reaches the site. Configure a traffic-acquisition view for AI referral sources actually present in your data, review landing pages, and check the conversion events that matter to the business. Do not assume that every Copilot citation produces a visible referrer. A person may open a cited page, return later through a branded search, paste a link into another browser, or never click at all.

Microsoft Clarity’s AI Visibility Citation dashboard is a separate tool worth understanding if the site uses Clarity. Microsoft says a project administrator can verify a domain by installing Clarity tracking code or connecting Google Search Console or Bing Webmaster Tools. Microsoft’s Clarity Citation dashboard documentation defines AI referral traffic as AI-referred sessions divided by total sessions and distinguishes it from page citations. Microsoft’s Clarity documentation is clear that citations do not measure traditional rankings, impressions, or click-through rates.

This is the right order of interpretation:

  1. A page was cited: source selection occurred in the reported AI surface.
  2. A measurable referral session arrived: a person followed a detectable route to the site.
  3. The session engaged or converted: the visit completed a defined on-site event.
  4. The business decided whether the change is material: the event had enough value, volume, and consistency to justify more work.

Do not skip from step one to step four. There may be influence without a measured referral, but that is a reason to describe an attribution gap, not a reason to assign revenue to citations.

Step 4: Use Google data as a parallel system, not a proxy for Bing

Google’s AI features are a useful comparison because Google documents a different reporting model. Google says AI Overviews and AI Mode traffic is included in Search Console’s Performance report under the Web search type, rather than in a separate universal citation report. It also states that no special technical requirement or special schema is required to appear, although normal eligibility, indexability, and SEO practices still apply. Google’s AI features documentation explains both the inclusion rules and the measurement boundary.

That does not make Search Console a Copilot dashboard. It means a team should preserve two separate questions: “What did Bing report as a citation across supported AI experiences?” and “What did Google report as Search performance, including its own AI features?” The interface, inventory, and aggregation differ.

The same caution applies when comparing Search Console with Google Analytics. Google describes Search Console as the source of truth for Google Search performance and Analytics as the source of truth for behaviour on the site. It also explains why clicks and sessions do not necessarily match. Google’s guide to Search Console and Analytics is a good model for the Bing-and-analytics comparison: use each product for the event it directly observes.

Turn citation data into page decisions

Measurement should lead to a content decision, not merely a more polished AI dashboard.

A cited page with weak on-site outcomes

Start by checking relevance rather than rewriting the page. Is the page a useful answer to the grounding language? Does its opening state the answer plainly? Are time-sensitive claims dated and sourced? Is the page technically accessible and linked from a relevant hub? A citation may expose an informational mismatch: the page is good enough to be used as a source but poor at serving the person who lands there.

Improve the page only where you can name the gap. For a technical guide, that may mean correcting a version-specific instruction and placing the limitation next to the instruction. For a comparison page, it may mean defining the comparison criteria and showing which option fits each constraint. This is the same editorial discipline behind citable content patterns: self-contained answers, named sources, and clear limits. It is not a promise that a sentence shape will win a citation.

A relevant page that is indexed but rarely cited

Microsoft recommends reviewing clarity, structure, completeness, evidence, freshness, and consistency across text, images, and video when using AI Performance insights. Those are reasonable audit prompts, not a checklist that guarantees selection. Microsoft lists these content-improvement directions in its AI Performance guidance.

Run a narrow audit before adding new sections:

  • Confirm that the page is the canonical, indexable version and has no competing near-duplicate page for the same intent.
  • Check whether the title, H1, opening answer, and headings agree on the question being answered.
  • Verify every material fact that could have aged, then update the visible source link rather than only changing the publication date.
  • Make the important explanation available as text, not only inside an image, video, or downloadable document.
  • Add a relevant internal link from a strong hub or companion page when it helps a reader move to the evidence or next task.

The technical SEO guide is a sensible place to revisit crawlability, indexability, canonicals, and internal discovery before treating AI visibility as a writing problem. If the underlying page cannot be reliably found or rendered, a clever paragraph is beside the point.

A citation increase after an edit

Record the edit, publish date, index-discovery action, and the reporting periods before and after it. Then wait for a comparable observation window. If citations increase, write “citation activity increased after the edit” rather than “the edit caused Copilot to cite us.” Other changes may have occurred in the engine, in the query mix, or in the content set.

This may sound overly careful. It is cheaper than spreading a false playbook across a content team. If the effect repeats over several updates and comparable pages, you have a stronger operational pattern. You still have a pattern, not an algorithm disclosure.

A page cited for outdated information

This deserves a fast response. Correct the claim, preserve or explain material changes where readers need the context, and make the new version discoverable. Microsoft says IndexNow notifies participating search engines when URLs are added, updated, or deleted; its AI Performance guidance describes it as a way to help current content be discovered across search and AI experiences. Microsoft’s IndexNow guidance in the AI Performance announcement supports notification, not a guaranteed recrawl, citation, or answer update.

The next useful action is to verify the live page, inspect the canonical and robots directives, submit the update through the site’s established discovery workflow, and monitor the cited URL in the next review period. Do not publish a new duplicate just to replace an old citation target. Consolidate the answer around the URL that should remain authoritative.

A compact monthly scorecard

Use a page-level scorecard rather than a single “Copilot performance” number. The following template keeps incompatible events separate while giving the team a decision surface:

URL AI citations Grounding-query sample Bing clicks / impressions AI referral sessions Primary conversion Editorial decision
/guide-a/ Record from AI Performance Record the associated phrases Record from Search Performance Record from analytics Record the chosen event Keep, refresh, expand, consolidate, or investigate

For each row, add a one-sentence interpretation with its evidence boundary. Examples:

  • “This page was repeatedly cited for implementation phrasing, but the lead paragraph does not answer that phrasing; revise the opening and verify the documentation link.”
  • “This page received citations and referral sessions, but its download conversion is flat; review the landing-page offer before producing more adjacent content.”
  • “This page has Bing search impressions but no reported citations; keep monitoring and audit duplicate intent before assuming an AI-content issue.”

The row should lead to one reversible action. Avoid a queue of generic tasks such as “improve authority” or “optimise for AI.” A decision is useful when another person can review the URL, the data window, the proposed change, and the reason for it.

Limits to put in every report

Copilot citation data is a measurement layer, not an attribution system. State these limits in the report so no one mistakes a dashboard for a causal model.

First, AI Performance is a public preview. Microsoft may revise coverage, metric definitions, or supported experiences. Preserve export dates and source documentation links with the report.

Second, a citation shows that a source was displayed. It does not state whether the source was prominent, whether the user read it, or whether the citation changed a purchase decision. Microsoft’s metric documentation says this directly.

Third, grounding queries are sampled retrieval phrases. They can help explain how a page was found, but they are not an exhaustive list of user prompts or a substitute for demand data.

Fourth, search, citation, referral, and conversion tools observe different events. Analytics implementation, consent choices, cross-device behaviour, and referrer handling can all affect what a team can see after a click. Google makes the same general point when it explains why its Search Console clicks and Analytics sessions do not match exactly. Use the documented discrepancy model instead of forcing the datasets to reconcile.

Finally, citations should never be used as a reason to lower editorial standards. Microsoft advises users to review cited sources and verify important details; its Bing search documentation describes source links as a way for people to check generated answers. Microsoft’s Bing search guidance supports that user behaviour. The publisher’s job is to make the underlying page accurate enough that a reader who clicks through finds the evidence, context, and limits intact.

The next useful action

Choose five URLs that already matter to the business. Capture 28 complete days of AI Performance data for them, identify cited pages and their grounding-query samples, then place the same URLs beside Bing Search Performance and analytics. Do not change every page at once. Pick one documented content gap, correct it, annotate the date, and review a comparable period.

That process cannot guarantee a Copilot citation. It produces an auditable record of what Bing reports, what visitors do when they arrive, and which content decisions have evidence behind them.

Share this article

If you found this content useful, share it with your colleagues.

Frequently Asked Questions

What does Bing AI Performance measure?

Bing AI Performance reports total citations, average cited pages, grounding queries, page-level citation activity, and citation trends across its supported AI experiences. Microsoft describes it as a public preview, so its coverage and metrics can change.

Does a Copilot citation mean my page ranks first?

No. Page-level citation activity shows how often a URL was cited, not its ranking or placement in an answer. Separately, Average Cited Pages is the daily average number of unique URLs, aggregated across supported AI surfaces, and does not indicate authority or the role of a page in an individual answer. Treat either metric as evidence that content was displayed as a source during the selected period.

Can Bing Webmaster Tools show Copilot clicks and impressions?

Bing Search Performance reports clicks and impressions from several Bing sources, including Chat. AI Performance focuses on citations. Keep the reports separate because they describe different events and may cover different AI surfaces.

What are grounding queries in AI Performance?

Grounding queries are key phrases used to retrieve content that was referenced in AI-generated answers. Microsoft says this data is a sample of overall citation activity, so it should guide investigation rather than be treated as a complete list of user searches.

Should I optimise a page only because it receives Copilot citations?

No. First check that the page is accurate, indexable, and useful for the associated retrieval intent. Then compare citations with Bing search data, referral sessions, and conversions before deciding whether to expand, consolidate, refresh, or leave the page alone.

Stay updated

Receive the latest articles, tips and strategies about SEO, web performance and digital marketing in your email.

We send a newsletter every week, and you can unsubscribe at any time.

Tags: #Bing Webmaster Tools #Microsoft Copilot #AI Performance #Copilot citations #GEO #AI visibility #Bing SEO
EG

Elu Gonzalez

SEO Expert & Web Optimization