Measurement

Why a single AI visibility score is noise

We ran twelve money queries five times each against one model on one day. Seven of them changed their answer. Here is what that does to every before-and-after published in this category.

Citon4 min read

The short version

A single AI visibility reading cannot be trusted, because the same query asked twice returns a different answer. In our own step-zero run, twelve queries repeated five times each on one model produced seven that flipped outcome across identical repeats. That is a property of the system, not a vendor defect. It does not mean citations are unmeasurable: the noise is unbiased, so it cancels between a treated group and a held-out control, which is why a difference between two arms can be trusted when a single level cannot.

If you have been sold an AI visibility score, you have been sold a single reading of a system that does not return the same answer twice. The number is not wrong exactly. It is just not a number in the way you are being invited to believe it is.

We tested this on ourselves before building anything on top of it, because the whole product depends on the answer.

The five minute test5 repeats

best rate limiting api

chatgpt.com
perplexity.ai
Named in 3 of 10 runsSame question

The test we ran, and what came back

Twelve money queries. Five identical repeats each. One model, one day, nothing changed between runs. Sixty calls, all sixty succeeded.

Seven of the twelve queries changed their outcome across those identical repeats.

Not changed wording. Changed outcome: named in one run, absent in the next, with the same question asked the same way minutes apart. If a vendor had run one of those queries once and shown you the result, you would have had a fifty-fifty chance of seeing the opposite finding.

This is not a defect in any particular model and it is not a vendor cutting corners. Retrieval is stochastic, the index moves, and the answer is generated rather than looked up. Instability is the system working as designed.

What that does to a before-and-after

Take the standard shape of a case study in this category. Measure a score. Do some work. Measure again. Report the difference.

If a single reading can flip on its own, then the difference between two readings contains the work you did plus however much the system moved by itself. Nothing in that method separates the two. A vendor showing you a twenty-point improvement has not shown you that they caused twenty points. They have shown you two draws from a distribution.

The uncomfortable version: run the same before-and-after with no work done at all, and you will still get a number. Sometimes a flattering one.

One query, five identical runsSample

one money query, nothing changed between runs

Named in 3 of 52 flips

7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.

Why this does not mean citations are unmeasurable

The category mostly concluded that AI citations cannot be measured, and stopped there. That inference is wrong, and the reason it is wrong is the entire basis for what we do.

The noise is unbiased. It pushes readings up as often as it pushes them down, and it does not prefer one set of queries over another. So it cancels.

We tested that too rather than assuming it. Across 20,000 random splits of our query set with no intervention applied, the difference between the two halves centred on zero: mean plus or minus 0.0016, standard deviation 0.215. Under the null hypothesis, the difference between two arms is genuinely zero even though each arm individually is jumping around.

That is the whole trick. You cannot trust a level. You can trust a difference between two groups measured the same way at the same time, because whatever the system did to one arm, it did to the other.

Causal lift+24pp
Treated ControlDay 0 to 90

What an honest reading actually costs

Once you accept you need two arms, the sample size stops being a detail and becomes the constraint.

Our twelve by five design has a minimum detectable lift of 51.1 percentage points. That is not a measurement instrument. A real engagement moves the number by five to fifteen points, so a design that can only see a fifty-one point move would report almost every genuine win as nothing.

Getting the floor down to 9.9 points takes forty queries by forty samples, which is about 3,200 calls per timepoint and roughly six hours of machine time. Twice, because you need a before and an after.

That cost is why almost nobody does it. It is also why a number produced this way means something.

What to ask a vendor

Three questions, and they are cheap to ask.

  • How many times did you sample each query, and what did the repeats disagree on?
  • What is your minimum detectable effect at that sample size?
  • What did the queries you did not work on do over the same period?

The third one is the one that matters. Without a held-out control, there is no answer to "compared to what", and every reported improvement is a level, not a lift.

The honest state of our own numbers

We have the instrument and it has passed its kill test. What we do not yet have is a completed causal-lift result, because that takes a full engagement and ninety days.

So we are not going to show you a lift number. We would rather show you the method and let you check it than show you a figure we cannot defend. If we could already show the after, we would not need a control group, and neither would anyone else.

Sources

Every number above, and where it came from. A figure without a row here is one we should not have printed.

Step-zero measurement run
12 money queries x 5 identical repeats, 60 of 60 calls succeeded, OpenAI gpt-5-search-api, single model, single day. Our own measurement.
Permutation test
20,000 random splits with no intervention applied. Null difference centred on zero: mean +0.0016 / -0.0016, SD 0.215.
Power analysis
The 12x5 design has a minimum detectable lift of 51.1 percentage points, roughly 4x underpowered. A 40 query x 40 sample design reaches 9.9pp.
Ahrefs, 43,000 keywords
AI Overviews persist about 2.15 days and 45.5 percent of citations change between consecutive observations, while the answer stays roughly 95 percent semantically identical.