Ask ChatGPT today which webshops are best in your category. Ask again tomorrow. There is a real chance the answer differs without anything on your site having changed. That is not a fault and not a bad prompt. It is how these models work.
The question that matters is: how bad is it? There is research on that now, and the answer is more uncomfortable than most tools in this market let on.
A variance analysis of AI answers about brands breaks down where the difference between two answers comes from:
| Source of variance | Share |
|---|---|
| Asking the same thing again | 34.8% |
| Language of the question | 31.6% |
| Brand combined with context | 29.6% |
| Which model you ask | 1.7% |
| How you phrase the question | 0.3% |
| The brand itself | 0.7% |
That last row is the point. The thing you want to measure, whether your brand actually matters to the answer, accounts for less than one percent of the difference between two answers. Everything else is context and chance.
The same research works out how best to spend a fixed measurement budget. Gain per extra measurement, relative to simply asking the same question again:
| Where the budget goes | Effectiveness |
|---|---|
| The same question again | 1× |
| A different phrasing | 4× |
| One more model | 5× |
| One more language | 15× |
This is counter-intuitive. The reflex on an unreliable measurement is to measure more often. But asking the same thing again yields almost nothing, while asking it in another language yields fifteen times as much. For a Dutch business that is not a trick either: your customers ask in both languages too.
If a single measurement carries barely any signal, then a dashboard showing a bare percentage that moves a few points each month is not insight. It is an invitation to wrong conclusions. Someone sees 28% become 34%, concludes the new copy is working, and builds a plan on it.
The honest version: report a margin, and say plainly when the difference falls inside it. "No detectable change" is a valid result. Usually it is the correct one.
One more number from the same research, and this is one you rarely see in a sales pitch. Even a full design with eight languages, three models and fifteen phrasings reaches a reliability of about 0.36 for ranking brands. That is directional, not exact.
We do not think that is a reason not to measure. It is a reason not to lie about it. A trend across several months says something. A single figure with two decimals says less than it appears to.