That’s what it said

How we measure

Most of the numbers in this category are noise.

Here is how we get ours, in plain language. How many times we ask, what we do about the fact that assistants disagree with themselves, and the five things we tell you we cannot measure at all.

The problem

Ask the same question twice. Get two different answers.

This is not a glitch. It is how these systems work. Which means a single run tells you nothing, and a tool that reports a number from a single run is reporting a coin flip.

One real questionasked three timesNames: Rival A, Rival B, Rival Cyou do not appearNames: Rival A, you, Rival Dyou appear, secondNames: Rival B, Rival Eyou do not appearSame question.Three answers.Report any one of them and youhave told the customer a story.
Illustrative, drawn to show the mechanism. Real runs vary more than three.

What we do instead

We ask it five ways, not five times.

Repeating one sentence measures the sentence. Real people do not all phrase things the same way, so we ask each buying question five ways a real person might ask it. Spending the same budget on five different phrasings rather than on repeating one tells you more, because how people ask is most of what moves the answer.

ONE BUYING QUESTION“Who should I usefor this?”FIVE PHRASINGSbest option for...what do people use for...alternatives to...is X worth it for...how do I choose...ASKED ONCE EACH5observationsof one question a persona asksacross dozens of questionsONE FIGURE, WITH ITS RANGE40%the range we would report alongside itnarrow across many questions, wide on any one
Figures shown are illustrative. The width of that range is the honest part. Any single question carries a wide one, which is exactly why a tool reporting a move from 38% to 44% as a win is reporting nothing. The headline figure combines dozens of questions, so its range is narrower than any one of them.

Asking five ways rather than one means the variation in how people phrase things sits inside that range instead of being hidden outside it. Our ranges look wider than other tools’ as a result. That is the point. We report a range on the figures that combine many questions, and we do not put a percentage on a single question, because five observations cannot carry one honestly.

Benchmark

The same six questions, asked of the category.

Ask any vendor these six things before you buy from them, including us. The answers are more revealing than any feature list.

Question to ask a vendorWhat we found in the categoryOur answer
How many times is each question asked?Often once. Three in the product we reviewed.Once each, in five different phrasings. Five observations per question.
What happens to the wording?One sentence, repeated.Five ways a real person might ask it.
Is a range shown around the number?No range we could find.On every presence figure, always.
When does a number count as a change?When it moves. An arrow appears.Only when the ranges stop overlapping.
Are the raw answers kept?Not stated.In full, forever, so history survives our own bugs.
Is the sampling design published?Not that we could find.This page.

Middle column from our own review of a leading platform in August 2026, working from its live product. We have not named it, because we hold ourselves to checking a claim before we publish it and a screenshot of one account is not a survey of the category. If a vendor you are considering answers these six differently, that is the answer that matters.

The limits

Five things we cannot measure, and nor can anyone else.

We would rather you hear these from us now than discover them in month three. Two of them are the reason this page exists.

We do not have the questions real people type

Nobody outside the platforms does. Every prompt set in this category, ours included, is built rather than observed. Ours is built from what people actually ask in forums, reviews and your own sales calls, and we say plainly that it is built.

We cannot see when you were read but not named

An assistant can read your page, use it, and never cite you. That happens, it counts, and no tool on the market can measure it.

We cannot separate what a model learned from what it just read

We can test with search turned off, which tells us what the model already believes about you. We cannot give you a clean percentage split between the two.

We cannot tie this to revenue

Someone asks an assistant, hears your name, and searches for you a week later. That shows up in your analytics as branded search. Anyone who claims to attribute it is measuring something else.

We do not simulate every consumer surface

We collect from routes the platforms sanction. That means we miss some logged-in and account-tier behaviour. We chose that over running a proxy network, and we publish how far apart the two are.

One more, and it is the one people find hardest to hear: this will not send you much traffic. Under 1% of visits for most sites. Source: our own measurements and published operator data, 2026. The point is being one of the three names the assistant says, not clicks.

What we publish about our own accuracy

We collect from routes the platforms sanction, so a fair question is how far that sits from what you would see typing the question yourself. Every quarter we check a fixed set by hand and publish the gap.

If it is small, it retires the strongest objection to how we work. If it is large, you learn it here rather than from a spot check of your own. Either way it is a number no other vendor in this category publishes.

See what AI says about your brand