The differentiator here is honesty about measurement, not a bigger score. This page explains exactly what AnswerRadar checks, how, and where the method falls short.
01 / What we measure
We don't guess at AI visibility — we query the surfaces themselves and report what comes back.
| Surface | What it is |
|---|---|
| ChatGPT search (consumer surface, logged-out) | Accessed via a third-party data provider that queries ChatGPT's search-enabled answers the way a logged-out consumer would, without any account or chat history attached. |
| Perplexity | Queried directly through Perplexity's Sonar API, which returns search-grounded answers with citations. |
| Gemini (consumer surface, logged-out) | Accessed via a third-party data provider that queries Gemini's consumer-facing answers the way a logged-out user would, without any account or chat history attached. |
02 / Why every prompt runs 5 times
Ask an AI model the same question twice and you can get two different answers — different sources cited, different brands named, sometimes a mention that appears in one run and vanishes in the next. That's not a bug in these tools; it's how sampling-based language models behave. A tool that runs your prompt once and hands you a score is showing you a single dice roll and calling it the truth.
So for every prompt, on every engine, we run it 5 times (N=5) and report the mention rate — how many of those 5 runs mentioned your brand — along with a Wilson 95% confidence interval around that rate. The interval tells you the plausible range the true mention rate falls in, given the sample size, instead of pretending 5 runs pin it down exactly.
A single-run tool would have shown you any one point in this range.
03 / What we can't measure
We disclose the surface instead of pretending otherwise. Specifically, we cannot measure:
API and scraped surfaces differ from what a logged-in user sees in their own app. We disclose the surface instead of pretending otherwise.
04 / Industry standard
"Citing what a single AI tool returns in response to a single prompt, at a single moment, in a single market, is methodologically weak evidence." — AMEC (International Association for the Measurement and Evaluation of Communication)
That's the exact failure mode repeated, multi-engine measurement with disclosed confidence intervals is built to avoid.
05 / Monthly calibration
Once a month, we manually run the same questions in the real consumer apps (logged in, as an ordinary user would) and publish how often that manual check agrees with our automated measurements.
| Period | Agreement rate |
|---|---|
| First calibration | launch week — results will be posted here. |
06 / Changelog
2026-07: Gemini measurement switched from Gemini API grounding to the consumer surface (better fidelity, disclosed).
2026-07: v1 methodology (5 prompts × 5 runs × 3 engines).