Two dashboards can show 10% and 25% for the same set of answers and both be mathematically correct. One might divide owned citations by all citations. The other might count how many citation-bearing answers included the brand's site.
Before asking which number is better, ask what was counted, what it was divided by and which observations were eligible. Those decisions determine what the result can tell you.
All numerical examples below are synthetic arithmetic exercises. They are not Beacon results, provider tests or customer benchmarks.
One fictional dataset produces several valid percentages
Imagine a collection job with these invented totals:
- 100 attempts, of which 80 produced valid answers and 20 failed
- 20 valid answers mentioned the resolved brand
- Eight valid answers recommended it for the requested job
- 40 valid answers contained at least one citation
- Ten valid answers contained at least one owned-domain citation
- 160 citation instances appeared in total, including 16 owned-domain instances
An “instance” here means one counted citation occurrence under a stated deduplication rule. A tool that collapses repeated URLs would produce a different count.
| Synthetic metric | Calculation | Result | What it describes |
|---|---|---|---|
| Collection completion | 80 / 100 | 80% | How much of the planned collection succeeded |
| Brand mention rate | 20 / 80 | 25% | Brand presence among valid answers |
| Recommendation rate | 8 / 80 | 10% | Recommendations among valid answers |
| Owned-citation presence | 10 / 80 | 12.5% | Valid answers with at least one owned link |
| Owned presence among citing answers | 10 / 40 | 25% | Owned links within the subset that cited anything |
| Owned citation-instance share | 16 / 160 | 10% | The owned portion of counted citation instances |
These measures answer different questions. “We have 25% AI visibility” discards the distinction. A useful report names the metric, shows the counts and identifies the population.
It should also specify domain ownership. Decide whether a documentation subdomain, regional store, third-party marketplace page or acquired brand belongs in the owned set. Preserve changes to that mapping.
Collection failures can create a false trend
Consider another synthetic comparison:
| Collection period | Attempts | Valid answers | Brand mentions | Mentions / valid answers | Mentions / attempts |
|---|---|---|---|---|---|
| A | 100 | 80 | 20 | 25% | 20% |
| B | 100 | 100 | 25 | 25% | 25% |
The mention rate among valid answers remains 25%. Dividing mentions by attempts makes the second period look five percentage points better because collection coverage improved.
The attempt-based measure is computable, but its meaning is different. Label it explicitly if you use it. Do not quietly assign failed or blocked attempts an absent-brand answer. Also investigate whether failures cluster around a platform, region or prompt type; the successful subset may be biased.
Retain separate statuses for a collected answer that contains no citations, an answer that omits the brand, a refused answer and a collection failure. A blank field cannot tell you which occurred.
Consumer search API and connected answers need separate tracks
A consumer assistant may retrieve sources, apply product settings or use account context. An API test has the tools and settings supplied by the test. A connected experience may already have the brand explicitly in play.
The following second exercise is also entirely invented:
| Synthetic track | Valid answers | Brand mentions | Rate | Suitable interpretation |
|---|---|---|---|---|
| Consumer web-search product | 20 | 8 | 40% | Presence in this consumer prompt panel |
| API with browsing disabled | 20 | 4 | 20% | Presence under this API configuration |
| Explicitly invoked brand connection | 20 | 20 | 100% | Brand presence after that connection was invoked |
Adding the rows produces 32 mentions in 60 answers, or 53.3%. That blended rate describes the constructed mix. It cannot stand in for the consumer-search rate, and it says nothing about how often a buyer would discover and invoke the connection.
Evaluate the connected track on its own job: correct product retrieval, fresh data, successful actions and appropriate handling of missing information. Use unbranded consumer questions to investigate earned discovery.
A changing prompt mix can manufacture growth
Suppose a fictional system mentions a brand in 80% of branded questions and 10% of unbranded questions. Keep those rates fixed and change only the sample:
| Synthetic panel | Branded questions and mentions | Unbranded questions and mentions | Combined rate |
|---|---|---|---|
| A | 40 questions, 32 mentions | 60 questions, 6 mentions | 38 / 100 = 38% |
| B | 80 questions, 64 mentions | 20 questions, 2 mentions | 66 / 100 = 66% |
The headline rises by 28 percentage points. Performance within both question groups is unchanged.
Keep a stable core panel for trends, and show discoveries from new questions separately. When the panel changes, retain its version and explain the effect on comparability. Use customer-derived questions where available, and label model-generated candidates as research hypotheses rather than observed demand.
A retrieved source and a citation have different denominators
AirOps' March 2026 observational study reported that 15% of 548,534 retrieved pages were cited across its 15,000-prompt study. It also reported a 43.2% citation rate among pages ranked first in the compared Google results. That second figure is conditional on rank; it is not the share of all citations supplied by first-ranked pages. The study does not establish the effect of editing a Beacon page. Original AirOps study
This distinction is a useful reading habit. “Among first-ranked pages, how many were cited?” and “Among cited pages, how many ranked first?” reverse the denominator. One result cannot answer both questions.
Only calculate retrieval-to-citation rates when you have an appropriate retrieval trace. A source mentioned in an answer is not a complete retrieval list, and a server log entry does not establish that the page was considered for a particular answer. Missing trace data should remain unknown.
Citation share can improve while the business result worsens
A citation identifies a source reference. It does not establish that the answer recommended the product, represented it correctly or caused a sale. Inspect the sentence the citation is meant to support.
A brand could receive a link in an answer that advises against buying its product. It could be recommended with a wrong specification. A citation could lead to a useful page while referring to the wrong regional offer. Keep recommendation, factual accuracy, citation fidelity and qualified business outcomes as separate measurements.
Bing makes a related boundary explicit for its own citation-share metric: it is observational and does not represent traffic share, quality or ranking. A third-party metric with the same label may use a different population. Bing's metric definition
Review a chart before acting on it
Ask for these details beside any important trend:
| Review question | Why it matters |
|---|---|
| Did the exact prompts or their weights change? | The sample can move the result without any underlying improvement |
| Did the provider, model, browsing mode or account condition change? | The collection may now be measuring a different experience |
| Did completion or citation-bearing-answer coverage change? | Missing data can alter the denominator |
| Did the brand/entity mapping or citation deduplication change? | A parser change can look like a market change |
| Are repeat runs independent enough for the stated uncertainty? | Repeated answers to related prompts are not independent buyers |
| What happened to untreated comparison questions? | Platform-wide movement can be mistaken for a content effect |
| Can a reviewer open the answer and check its cited source? | A number without evidence cannot resolve a disputed judgment |
A practical export should retain the counts and definitions used to create the chart. Keep old definitions available when the method changes, and mark the break in the series.
Use the audit methodology to design collection and review. Use the product-data checklist when the answer reveals a factual conflict. Before approving the next content task, make sure the reported change survives these basic checks.