How we measure

The differentiator here is honesty about measurement, not a bigger score. This page explains exactly what AnswerRadar checks, how, and where the method falls short.

01 / What we measure

Three surfaces, queried directly

We don't guess at AI visibility — we query the surfaces themselves and report what comes back.

SurfaceWhat it is
ChatGPT search
(consumer surface, logged-out)
Accessed via a third-party data provider that queries ChatGPT's search-enabled answers the way a logged-out consumer would, without any account or chat history attached.
PerplexityQueried directly through Perplexity's Sonar API, which returns search-grounded answers with citations.
Gemini
(consumer surface, logged-out)
Accessed via a third-party data provider that queries Gemini's consumer-facing answers the way a logged-out user would, without any account or chat history attached.

02 / Why every prompt runs 5 times

A single AI answer is a dice roll

Ask an AI model the same question twice and you can get two different answers — different sources cited, different brands named, sometimes a mention that appears in one run and vanishes in the next. That's not a bug in these tools; it's how sampling-based language models behave. A tool that runs your prompt once and hands you a score is showing you a single dice roll and calling it the truth.

So for every prompt, on every engine, we run it 5 times (N=5) and report the mention rate — how many of those 5 runs mentioned your brand — along with a Wilson 95% confidence interval around that rate. The interval tells you the plausible range the true mention rate falls in, given the sample size, instead of pretending 5 runs pin it down exactly.

A single-run tool would have shown you any one point in this range.

03 / What we can't measure

Honest limits

We disclose the surface instead of pretending otherwise. Specifically, we cannot measure:

API and scraped surfaces differ from what a logged-in user sees in their own app. We disclose the surface instead of pretending otherwise.

04 / Industry standard

This isn't just our opinion

"Citing what a single AI tool returns in response to a single prompt, at a single moment, in a single market, is methodologically weak evidence." — AMEC (International Association for the Measurement and Evaluation of Communication)

That's the exact failure mode repeated, multi-engine measurement with disclosed confidence intervals is built to avoid.

05 / Monthly calibration

Checking our own work

Once a month, we manually run the same questions in the real consumer apps (logged in, as an ordinary user would) and publish how often that manual check agrees with our automated measurements.

PeriodAgreement rate
First calibrationlaunch week — results will be posted here.

06 / Changelog

Method history

2026-07: Gemini measurement switched from Gemini API grounding to the consumer surface (better fidelity, disclosed).

2026-07: v1 methodology (5 prompts × 5 runs × 3 engines).