How we measure

The differentiator here is honesty about measurement, not a bigger score. This page explains exactly what AnswerRadar checks, how, and where the method falls short.

01 / What we measure

Three surfaces, queried directly

We show you the answers to the questions exactly as they came back, rather than predicting what the AI would say.

SurfaceWhat it is
ChatGPT search(consumer surface, logged-out)Accessed via a third-party data provider that queries ChatGPT's search-enabled answers the way a logged-out consumer would, without any account or chat history attached.
Perplexity(consumer surface, logged-out)Accessed via a third-party data provider that queries Perplexity's consumer-facing answers the way a logged-out user would, without any account or search history attached.
Gemini(consumer surface, logged-out)Accessed via a third-party data provider that queries Gemini's consumer-facing answers the way a logged-out user would, without any account or chat history attached.
Google AI Mode(consumer surface, logged-out)The AI answer Google shows in its own search results, which is a different surface from Gemini and cites different pages. We measured both on the same question at the same moment and they shared almost no sources, so covering Gemini does not cover Google search.

02 / Why every prompt runs 7 times

A single AI answer is no different from a roll of the dice

Ask an AI model the same question twice and you can get two different answers. Different sources cited, different competitors named, sometimes a mention that shows up in one run and vanishes in the next. That isn't a bug in these tools; it's how sampling-based language models behave. A tool that runs your prompt once and hands you a score is showing you a single dice roll and calling it the truth.

So for every prompt, on every engine, we run it 7 times (N=7) and report the mention rate: how many of those 7 runs named your brand. We do not claim 7 runs measure the rate exactly; within the limits of that sample, we show a Wilson 95% confidence interval alongside the mention rate.

03 / What we can't measure

Honest limits

We state up front which surface we measure. Here is what we cannot measure:

  • Logged-in personalization. A signed-in user's account, browsing history, and settings can change what an AI assistant shows them. We measure logged-out or API-level responses.
  • Chat memory. Prior conversations can shape later answers for a returning user. Our queries carry no memory between runs.
  • Location effects. Answers can shift by country, city, or IP-inferred location. We run from a fixed vantage point, not from everywhere your customers are.
  • Silent model routing. Providers route requests to different underlying models without announcing it, so "ChatGPT" or "Gemini" on a given day may not be the exact model version we tested against.

API and scraped surfaces differ from what a logged-in user sees in their own app. We disclose the surface instead of pretending otherwise.

04 / Why one run is not evidence

Why a single measurement is uncertain

A mention in a single measurement is quite likely a one-off, and there is no guarantee the next run gives the same answer. AI writes its answers slightly differently each time, so a competitor named this time may not be named in the next measurement.

So we measure the same questions repeatedly, across several platforms, and show a confidence interval with the result. It can look less flashy than a service that shows one score, but it is what lets you tell noise from real change when the numbers move.

05 / Changelog

Method history

2026-08-21: added Google AI Mode as a fourth platform, on every plan including free scans. It is the AI answer inside Google's own search results, and our own measurement is the reason it is here: asked the same question at the same moment, Google AI Mode and Gemini shared two sources out of seventy-seven. It also cites places the others do not, including YouTube and Reddit. A scan is now 280 runs, and Pro scans are 700.

2026-08-21: Perplexity measurement switched from the Sonar API to the consumer surface, so every platform we report is now read the way a logged-out visitor sees it. Sources per answer fell from roughly 19 to roughly 10, because the consumer page lists fewer of them than the API returned. A trend line that crosses this date is comparing two different surfaces, and we would rather say so than quietly redraw it.

2026-08-08: repeats per prompt raised from 5 to 7, putting a standard scan at 210 runs. Pro scans cover 25 prompts, or 525 runs. Cited sources are now recorded from what each engine returns alongside its answer. Until this date we read them back out of the answer text instead, which captured the site but not the page, and could list a site the engine had not actually cited.

2026-08-06: prompts per scan raised from 5 to 10, which took a scan from 75 runs to 150. Per-question cells moved to raw fractions instead of percentages, because at five runs the confidence interval spans roughly ±35 points, which reads as precision we did not have.

2026-07: Gemini measurement switched from Gemini API grounding to the consumer surface (better fidelity, disclosed).

2026-07: v1 methodology (5 prompts × 5 runs × 3 engines).