Nederlands
← all articles

One measurement is noise, and that is not an opinion

Ask ChatGPT today which webshops are best in your category. Ask again tomorrow. There is a real chance the answer differs without anything on your site having changed. That is not a fault and not a bad prompt. It is how these models work.

The question that matters is: how bad is it? There is research on that now, and the answer is more uncomfortable than most tools in this market let on.

Brand identity explains 0.7% of the variance

A variance analysis of AI answers about brands breaks down where the difference between two answers comes from:

Source of varianceShare
Asking the same thing again34.8%
Language of the question31.6%
Brand combined with context29.6%
Which model you ask1.7%
How you phrase the question0.3%
The brand itself0.7%

That last row is the point. The thing you want to measure, whether your brand actually matters to the answer, accounts for less than one percent of the difference between two answers. Everything else is context and chance.

In practice: ask once and you have mostly measured noise, with a hint of brand in it.

Breadth beats repetition, by a factor of fifteen

The same research works out how best to spend a fixed measurement budget. Gain per extra measurement, relative to simply asking the same question again:

Where the budget goesEffectiveness
The same question again
A different phrasing
One more model
One more language15×

This is counter-intuitive. The reflex on an unreliable measurement is to measure more often. But asking the same thing again yields almost nothing, while asking it in another language yields fifteen times as much. For a Dutch business that is not a trick either: your customers ask in both languages too.

What this means for a report

If a single measurement carries barely any signal, then a dashboard showing a bare percentage that moves a few points each month is not insight. It is an invitation to wrong conclusions. Someone sees 28% become 34%, concludes the new copy is working, and builds a plan on it.

The honest version: report a margin, and say plainly when the difference falls inside it. "No detectable change" is a valid result. Usually it is the correct one.

We put our own first measurement next to this. Over twenty-five days, at a test domain, almost nothing moved: Claude named it equally often in both runs, Gemini likewise, ChatGPT went from zero-or-one to two out of eight. That last one looks like progress and probably is not: across eight questions the uncertainty runs from roughly 12% to 65%. Everything overlaps with everything.

The limit of what is measurable

One more number from the same research, and this is one you rarely see in a sales pitch. Even a full design with eight languages, three models and fifteen phrasings reaches a reliability of about 0.36 for ranking brands. That is directional, not exact.

We do not think that is a reason not to measure. It is a reason not to lie about it. A trend across several months says something. A single figure with two decimals says less than it appears to.