An AI answer audit should let another person check what was asked, what came back and why a factual claim was judged correct or wrong. A chart without those records gives a buyer little basis for deciding what to fix.
This page describes a practical audit standard for product brands. It is a proposed protocol, not a statement that every step is already implemented in Beacon. Product owners should confirm the collection, retention, validation and export capabilities before adopting it as an operating promise.
Define the question the audit needs to answer
Choose the decision before collecting responses. A brand might need to know whether buyers encounter it during product selection, whether answers describe its range correctly, or whether a connected stock lookup returns current information. Each requires different evidence.
Write a one-paragraph scope covering the category, market, language, intended buyer and decision. Name the platforms and surfaces you will test. Decide which competitors belong in the comparison before seeing the results.
Keep three types of observation separate:
- A consumer assistant searching the web
- A model or API answering without web retrieval
- An assistant using an explicitly connected or invoked product tool
Beacon's public FAQ already distinguishes browsing-enabled answers from model-memory answers. Preserve that distinction in every capture and report. Beacon FAQ
Build a buyer-question set that can be repeated
Start with questions from genuine sales, support and buying conversations that you are authorized to use. Remove personal information. Include comparisons, requirements, price constraints, availability and reasons to reject a product.
Keep discovery questions separate from branded questions. Asking an assistant about your named brand measures its description of that brand; it gives little evidence that the assistant would have introduced the brand on its own.
Version the question set. Assign each question an intent, market and reason for inclusion. Preserve the exact wording for the core trend panel. A second, clearly identified set of paraphrases can test whether a finding survives ordinary changes in how a buyer asks.
Do not describe the panel as the full market. It is a chosen sample of questions, and its composition will influence the result.
Preserve the conditions and the complete answer
For each attempt, record:
- Exact prompt and prompt-set version
- Time, provider, product surface and visible model/version
- Market, language, account condition and fresh-session status
- Browsing state and any connected tools or explicit invocation
- Complete answer, visible citations and a durable capture reference
- Collection status, including failures and unavailable responses
Use a fresh session and consistent settings where the product permits. A logged-in consumer interface and an API may expose different behavior. Label the collection channel; neither should stand in for the other without evidence.
Repeat questions in independent sessions on different days. Three repetitions per question and surface can be a practical pilot, but it is a starting design choice rather than a statistical guarantee. Increase the sample when answers vary or the decision has substantial consequences.
A failed collection is missing data. Preserve it in the ledger instead of counting it as an answer that omitted the brand. Follow provider terms and access controls; do not bypass a blocked service to complete the sample.
Resolve the brand before counting a mention
Brand names are not always unique. Resolve the answer to the correct company using a cited domain, owner, product description or other clear context. A reference to another product with the same name should not increase your visibility score.
Record unresolved cases for review. This matters for Beacon: the audited product is Beacon by EVAA at beacon.evaa.sg. A name match by itself is insufficient.
Next distinguish a mention, a substantive recommendation and a citation. “Other options include…” is a different buying signal from a recommendation tied to the user's requirements. A link to the brand's page is a separate observation again.
Check claims against a versioned product record
Prepare the reference record with the product-data readiness checklist before assessing answers against it.
Break the answer into factual claims a reviewer can assess. Keep the original sentence so the claim is not detached from a qualification.
A review record should identify the product and variant, market, attribute, answer value, source value, source URL or approved catalogue record, effective date and judgment. Useful judgment categories are:
| Judgment | Meaning |
|---|---|
| Supported | The appropriate source supports the claim with the same scope |
| Incorrect | The claim conflicts with the appropriate source |
| Outdated | It matches an earlier version but no longer applies |
| Unsupported | The available evidence does not substantiate it |
| Ambiguous | The product, market, date or meaning cannot be resolved |
| Unassessable | The audit lacks an appropriate reference |
An absence of evidence does not establish that a statement is false. Preserve those separate categories.
Catalogue data also has limits. A brand can usually supply its published dimensions, warranty or price history. A statement about comfort or long-term reliability may need testing or independent evidence. Do not mark every favorable subjective claim correct because it matches marketing copy.
Check whether the citation supports the claim
Open the cited page, verify its relevance and preserve the passage needed for review within applicable rights and retention rules. Check the product version, date and market.
A working link can point to a source that never makes the attached claim. Record citation fidelity separately from factual accuracy. An answer may be correct for a reason its citation does not establish, or cite a relevant page while misreading it.
Automated extraction and model judging can help reviewers find issues. Humans should adjudicate material errors and ambiguous cases, and check a sample of automated judgments. Record the judging procedure so a change of evaluator does not silently change the trend.
Publish the numerator and denominator
Use a small set of measures that answer distinct questions:
| Measure | Definition |
|---|---|
| Mention rate | Valid unbranded answers mentioning the resolved brand, divided by valid unbranded answers |
| Recommendation rate | Valid unbranded answers recommending the brand for the requested job, divided by valid unbranded answers |
| Domain citation rate | Valid answers citing the specified domain, divided by valid answers |
| Factual support rate | Supported claims, divided by assessed checkable claims, with unresolved categories disclosed |
| Citation fidelity | Supported claim–citation pairs, divided by checked pairs |
| Panel share of voice | Brand mentions, divided by mentions of the predefined competitor set in the panel |
The citation-metric examples show how different denominators and a changing prompt mix alter the result.
Show counts beside percentages. If unresolved claims are excluded from a denominator, say how many were excluded. Keep separate charts for different markets and collection modes unless you disclose and justify the weighting.
Sentiment does not measure correctness. A positive answer can contain the wrong specification. A lower mention rate can coexist with more accurate recommendations.
Combine sampled answers with observed site data
Use first-party tools to answer questions the prompt panel cannot. Google's current Generative AI performance report covers impressions from AI Overviews and AI Mode, with page, country, date and device views. Check actual property availability; it does not provide the factual review described here. Search Console report
Bing's AI reporting provides citation observations across supported Microsoft surfaces. Its citation-share measure is distinct from traffic share or a ranking score. Bing reporting
For commercial value, track qualified enquiries and agreed downstream events. OpenAI documents a ChatGPT referral parameter that can help identify visits, but referral data alone cannot capture every influence on a purchase. OpenAI publisher FAQ
Test a correction and document what happened
Tie each proposed fix to an observed problem. For a stale specification, correct the authoritative record and inconsistent public pages. For an unavailable connected action, repair and test the supported flow. Record what changed and when.
After confirming that the change is live, rerun the same prompt panel under comparable conditions. Where possible, retain an unchanged comparison group. Note provider changes, seasonality and other work that could affect the result.
Describe an observed increase as an observed increase. A before-and-after chart by itself cannot establish that the fix caused it. Avoid converting a sampled mention into a claimed impression, exact buyer count or media-equivalent revenue.
What the report should contain
Deliver the scope, prompt manifest, capture ledger, claim reviews, metric definitions, limitations and prioritized fixes. Include enough evidence for a client to dispute a judgment and for another operator to repeat the test.
Before buying an audit, ask to inspect one complete record. It should take you from a buyer's question to the answer, the cited source, the judgment and the proposed fix. The vendor decision guide applies that test to different kinds of tooling and implementation work.
For the published commercial units, review Beacon monitoring plans, then confirm which parts of this proposed protocol the agreed service can deliver.