Three questions, two similarly named products, and one source disagreement. This small Beacon investigation shows why a useful AI answer audit needs more than a correct-or-incorrect score.
On 9 October 2026, we asked ChatGPT about two UK Philips Hue Go variants. The assessed brightness and IP-rating values matched the official product records. One answer nevertheless began with a misleading “Yes”. Another appeared to conflict with the source until we checked the rendered store page.
This is independent public-product research, not a Philips Hue or Signify client case study. It does not establish an assistant accuracy rate, a successful correction or a commercial outcome.
What we tested
We fixed three questions before collecting answers, then ran each once in a fresh, signed-out conversation on the ChatGPT consumer website. We attached no plugin and supplied no Beacon context or source URL. The questions requested UK product information; the actual connection location and exact model were not disclosed. All conversations used the same browser.
The answer captures were taken on 8 October at 16:43 UTC, which is 9 October at 00:43 in Singapore. We retained the questions, answer text, source drawers and completion receipts. The public evidence file includes the complete captured transcript text and cited URLs, with account/session identifiers omitted.
Download the prompts, answers and source observations (JSON).
| Check | Exact question |
|---|---|
| P1: ambiguous name | For a UK buyer, how bright is the Philips Hue Go, and can I use it outdoors? |
| P2: exact product | For the white Philips Hue Go portable table lamp, EAN 8719514404571, what are its lumen outputs at 2700K and 4000K, its IP rating, and its current availability from Philips Hue UK? |
| P3: false premise | Is the Philips Hue Go portable accent light 7602031PU rated IP54? What are its lumen outputs at 2700K and 4000K for the UK model? |
Identify the product before judging its answer
The two UK pages we read had matching brightness figures but different IP ratings. Treating “Hue Go” as a single unambiguous product would lose that distinction.
| Product record | Identifier | At 2700K | At 4000K | IP rating |
|---|---|---|---|---|
| White portable table lamp | EAN 8719514404571 | 370 lm | 530 lm | IP54 |
| Portable accent light | SKU 7602031PU; EAN 8718696174005 | 370 lm | 530 lm | IP20 |
These are dated observations from the official UK table-lamp page and accent-light page. They are not a current purchasing or outdoor-safety recommendation. We did not assess battery-runtime or broader use claims.
P1 handled the ambiguity by conditioning its main answer on the table lamp and separately describing the accent light as IP20. The assessed numbers matched our source records. The retained source drawer listed the table-lamp page, however, without a separate accent-light source. A fact can match an independently checked record while the answer's retained citations provide incomplete support.
A correct specification can follow a misleading opening
P3 asked whether the exact accent-light SKU was IP54. Its opening was:
Yes—but there’s an important model distinction.
The same answer then stated:
IP rating: IP20, not IP54.
The later rating and temperature-specific lumen values matched the official record for that SKU. The opening still answered the yes/no question in the wrong direction. We classified this as a mixed answer: correct assessed facts, misleading initial wording. Calling the whole answer correct would hide the problem; calling all its specifications wrong would misrepresent the evidence.
An editorial rewrite could begin: “No. For UK model 7602031PU, the listed rating is IP20. Its light output is 370 lm at 2700K and 530 lm at 4000K.” This is our suggested wording, not a later observed answer or a correction implemented by the manufacturer.
Check the source before declaring an AI error
P2 identified the exact table-lamp EAN, returned the matching lumen values and IP54 rating, and said the item was in stock. Our web extraction of the cited official page had returned “Item no longer available”. That looked like an answer/source mismatch.
We then opened the official UK product page in the browser. At 16:43:37 UTC, its rendered view showed “In stock” for the same EAN. That agreed with the assistant.
| Observation | Availability shown | What it establishes |
|---|---|---|
| Official-page web extraction, study date | Item no longer available | The extracted representation disagreed with the answer |
| P2 answer, captured 16:43:15 UTC | In stock | What the assistant told the user |
| Rendered official page, captured 16:43:37 UTC | In stock | The browser view agreed with the answer |
We retained both source observations and classified availability as source-view disagreement, not a confirmed AI error. We did not test checkout stock or establish why the views differed. Rendering, caching, timing and locale are possible explanations, not findings. A store label also cannot establish global discontinuation.
This changed the audit conclusion. An automated check against just the extracted text would have flagged an error that our evidence could not substantiate.
Keep fact, citation and clarity separate
The useful output from this sample is a short evidence ledger:
| Run | Assessed facts | Citation support and clarity |
|---|---|---|
| P1 | Variant distinction and specified values matched | Retained drawer supported the table lamp; separate accent-light citation absent |
| P2 | Lumen values and IP rating matched | Exact product page cited; availability agreed with its rendered view but source representations conflicted |
| P3 | Later IP20 and lumen values matched | Exact accent-light page cited; opening “Yes” contradicted the later answer |
We did not verify every extra statement in the answers. P2's additional Signify document, P3's extra fixture-output claim, and broad safety, battery and suitability advice remain outside this bounded assessment. The downloadable transcripts retain them so a reader can see the limits of our review.
Three one-off answers from one surface cannot estimate typical performance. We did not run another assistant, repeat the prompts, change the manufacturer's data or measure an improvement. This is a demonstration of the checking method, not a benchmark.
What your team can take from the check
For a product-answer audit, ask for the exact question, dated answer, variant and market, source record, cited URL and a reason for each judgment. Keep unresolved source conflicts visible. Our audit methodology and product-data checklist explain how to scope those records for your own catalogue.
For a connected application, the same discipline helps define the data and actions an assistant should receive: exact product or record IDs, account permissions, current values and explicit limits. A connected build creates a controlled workflow where your integration is used; it does not establish that unrelated public answers will change. Our ChatGPT app and plugin offering is a separately scoped implementation service.
Bring one recurring customer question and the record your team considers authoritative. We can use those to scope an audit, decide whether repeat monitoring is useful, or assess a connected workflow. This study's next research step is a later repeat with the same questions and newly dated source observations.