New: the nine ways developer tools go invisible in AI answers.
citon
Measurement

Why a single AI visibility score is noise

Twelve money queries, five identical repeats, one model, one day. Seven changed their answer, which breaks every before-and-after published in this category.

Citon24 min read

Why a single AI visibility score is noise. 12 money queries, 5 identical repeats each, 7 flipped outcome.

Short answer

Can a single AI visibility score be trusted?

A single AI visibility reading cannot be trusted, because the same query asked twice returns a different answer. In our own step-zero run, twelve queries repeated five times each on one model produced seven that flipped outcome across identical repeats. That is a property of the system, not a vendor defect. It does not mean citations are unmeasurable. Measuring a DIFFERENCE between a treated group and a held-out control, rather than a single level, is the standard way to handle a noisy metric, because movement shared by both arms subtracts out. Two caveats we state rather than bury: symmetry across arms is an assumption borrowed from controlled-experiment design, not something anyone has measured on citation data, and it holds only while both arms would have moved in parallel anyway.

If you have been sold an AI visibility score, you have been sold a single reading of a system that does not return the same answer twice. The number is not wrong exactly. It is just not a number in the way you are being invited to believe it is.

We tested this on ourselves before building anything on top of it, because the whole product depends on the answer. The method that came out of it is the one we run on every answer engine optimization engagement, and the arithmetic behind it is written up separately in how we measure lift.

Seven of twelve queries flipped outcome across five identical repeats.
The step-zero run. Twelve money queries, five identical repeats each, one model on one day. Sixty of sixty calls succeeded, and seven of the twelve questions changed their answer between repeats.

At a glance

What a single score hides

AxisWhat a single score hides
EnginesA score averages ChatGPT, Perplexity and AI Overviews, which disagree on the same query more often than they agree.
QueriesOne number over a query set you did not choose hides which of YOUR money questions you lose.
VolatilityAnswers move run to run. A score with no trial count cannot tell drift from noise.
CausalityA number that moves tells you nothing about what moved it, or whether you did.
The five minute test5 repeats

best rate limiting api

chatgpt.com
perplexity.ai
Named in 3 of 10 runsSame question

The test we ran, and what came back

Twelve money queries. Five identical repeats each. One model, one day, nothing changed between runs. Sixty calls, all sixty succeeded.

Seven of the twelve queries changed their outcome across those identical repeats.

Not changed wording. Changed outcome: named in one run, absent in the next, with the same question asked the same way minutes apart. If a vendor had run one of those queries once and shown you the result, you would have had a fifty-fifty chance of seeing the opposite finding.

This is not a defect in any particular model and it is not a vendor cutting corners. Retrieval is stochastic, the index moves, and the answer is generated rather than looked up. Instability is the system working as designed.

Nor is it only us seeing it. SparkToro, testing repeated identical prompts, put the chance of getting the same brand list twice at under one in a hundred. And a statistical framework for generative-search measurement found two competing brands whose citation share overlapped inside the margin of error, which is a dashboard reporting a win that was noise.

From the field

The instability is not ours alone, and it has been measured independently

Testing repeated identical prompts, SparkToro put the chance of two identical asks returning the same brand list at under one in a hundred. Separately, a statistical framework for generative-search measurement found two competing brands whose citation share overlapped inside the margin of error on a single sample, which is a dashboard reporting a win that was noise.

SparkToro and Gumshoe, repeated-prompt consistency

It also is not evenly distributed. The questions that move most are the ones with the most plausible answers, which for a developer tool or API is exactly the shortlist question you care about: the one where six vendors are all defensible and the model has no strong reason to prefer any of them.

What that does to a before-and-after

Take the standard shape of a case study in this category. Measure a score. Do some work. Measure again. Report the difference.

If a single reading can flip on its own, then the difference between two readings contains the work you did plus however much the system moved by itself. Nothing in that method separates the two. A vendor showing you a twenty-point improvement has not shown you that they caused twenty points. They have shown you two draws from a distribution.

Four steps of a before-and-after measurement, ending with a difference that mixes the work done with the system's own movement.
The standard before-and-after. Nothing in the method separates what you did from what the system did on its own, which is why the last step reports a level rather than a lift.

The uncomfortable version: run the same before-and-after with no work done at all, and you will still get a number. Sometimes a flattering one.

A score you cannot decompose is a score you cannot act on. If it drops, your next question is which engine, which query, and did we cause it, and the number cannot answer any of the three.
The measurement position
One query, five identical runsSample

one money query, nothing changed between runs

Named in 3 of 52 flips

7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.

Why this does not mean citations are unmeasurable

The category mostly concluded that AI citations cannot be measured, and stopped there. That inference is wrong, and the reason it is wrong is the entire basis for what we do.

The noise is unbiased, and this is the load-bearing assumption of everything that follows, so it is worth being exact about what is and is not established here. That noise EXISTS is measured, by us and by others: temperature zero is not deterministic, and the reason is batching and GPU kernel reduction order rather than anything a prompt controls. That the noise is symmetric and independent across a treated and a held-out arm of AI-visibility queries is the standard assumption of controlled-experiment design, and we have not found anyone who has tested it directly on citation data. We assume it, we tested it on our own query set, and we are telling you it is an assumption rather than a finding.

We tested that too rather than assuming it. Across 20,000 random splits of our query set with no intervention applied, the difference between the two halves centred on zero: mean plus or minus 0.0016, standard deviation 0.215. Under the null hypothesis, the difference between two arms is genuinely zero even though each arm individually is jumping around.

That is the whole trick. You cannot trust a level. You can trust a difference between two groups measured the same way at the same time, because whatever the system did to one arm, it did to the other.

This is also the reason the tactical guides in this category are not wrong so much as incomplete. The tactics are broadly the ordinary ones and Google says so itself; what changed is the instrument. That argument is laid out in full in LLM SEO, what it actually is and what the work looks like.

Causal lift+24pp
Treated ControlDay 0 to 90

What an honest reading actually costs

Once you accept you need two arms, the sample size stops being a detail and becomes the constraint.

Our twelve by five design has a minimum detectable lift of 51.1 percentage points. That is not a measurement instrument. A real engagement moves the number by five to fifteen points, so a design that can only see a fifty-one point move would report almost every genuine win as nothing.

What each design can actually detect

DesignCalls per timepointMinimum detectable liftVerdict
12 queries x 5 samples6051.1 pointsOur pilot. Cannot see a real engagement.
20 queries x 20 samples40021.4 pointsStill coarser than the effect being measured.
40 queries x 40 samples3,2009.9 pointsUsable. Roughly six hours of machine time, twice.
Computed from our own step-zero variance, not from a vendor table. A real engagement moves the number by five to fifteen points, which is why the first row reports almost every genuine win as nothing.

Getting the floor down to 9.9 points takes forty queries by forty samples, which is about 3,200 calls per timepoint and roughly six hours of machine time. Twice, because you need a before and an after.

That cost is why almost nobody does it. It is also why a number produced this way means something.

What the noise actually is

It helps to be specific about where the variation comes from, because "the model is random" is both true and useless for deciding what to do about it.

Three things move between two identical asks. Retrieval picks a slightly different set of documents, because ranking is not deterministic and the index behind it is being rewritten continuously. Generation samples, so even given identical retrieved context the wording and therefore the named entities can differ. And the answer is assembled rather than looked up, so a source that was retrieved is not guaranteed to be cited, and one that was cited last time may be dropped in favour of a shorter path to the same claim.

That third one is the one people miss. Being retrieved and being named are different events, and only the second is what you are being measured on. A page can gain ground in retrieval for weeks without moving a single visible citation, which looks like nothing happening right up until it looks like a step change.

Three questions to ask a vendor, each paired with what a real answer contains.
The three questions, and what an answer has to contain to count. The third is the one that matters, because without a held-out arm there is no answer to compared to what.

The two-arm design, in plain terms

A control group sounds heavier than it is. It is one decision made before you start.

Take the query set you actually care about and split it. Work one half. Leave the other half alone, deliberately, for the whole period. Measure both at the same time, with the same prompts, on the same day, at the same trial count. The comparison you report is the difference between the two halves, not the change in the worked half.

The reason this works is the property established above: the noise does not know which half you chose. Whatever the system did to the treated arm over those ninety days, it did to the held-out arm as well, so the shared movement subtracts out and what is left is attributable. This is difference in differences, and it is old.

It carries one condition that is easy to state and easy to forget: the two arms have to have been moving in parallel anyway. A model update or index refresh mid-window that lands harder on one arm's queries than the other's breaks that, and no amount of sampling repairs it. There is also a distinction worth keeping: drift is not noise. Per-call sampling variance cancels in a difference; a systematic behavioural change correlated with time does not, if it hits the arms unevenly.

STEPS

The two-arm design, start to finish

  1. Split the query set before anything else

    Day 0

    Take the questions you actually care about and divide them. The split is made before any work starts, so it cannot be drawn later around a result you already like.

  2. Measure both arms at once

    Day 0

    Same prompts, same day, same trial count, both halves. A baseline taken at two different times is two baselines.

  3. Work one arm only

    Days 1 to 90

    The held-out half is left alone on purpose for the whole period. This is the expensive part and it is the part that makes the result attributable.

  4. Re-measure both, identically

    Day 90

    Same instrument as day 0. Changing the prompts or the trial count between timepoints silently changes what the difference means.

  5. Report the difference between arms

    Day 90

    Not the change in the worked arm. Whatever the system did on its own over ninety days, it did to both halves, so it subtracts out.

What it costs you is real and worth stating plainly. You are choosing not to work a set of queries you could have worked, for a quarter. That is the price of being able to say you caused something, and it is why almost nobody pays it.

What counts as a citation

Everything above assumes we agree on the unit. We should not, because vendors do not, and it is the quietest way two visibility numbers become incomparable.

What we count as a citation, and what we refuse to

CaseCountsWhy
Named in the answer textYesIt is what the reader actually sees and acts on.
Listed only in a collapsed source panelNoMost readers never expand it, so it is presence without effect.
Retrieved into context but not namedNoBeing available to the model is not the same as being recommended by it.
Named in response to a query containing your brandNoThat is existing awareness measured back to you, not acquisition.
Named as a negative comparisonCounted separatelyAppearing as the option to avoid is not a win, and averaging it in hides that.
A narrower rule than most vendors use, which produces smaller numbers. Stated here so a comparison against a looser definition is visible rather than silent.

Our rule is deliberately narrow. A citation is your brand or product named in the answer text a user reads, in response to a query they would plausibly type. Not linked in a source list they have to expand. Not retrieved into context and then dropped. Not present because the query contained your name already.

That last exclusion matters more than it sounds. A branded query is a question from somebody who already knows you, and counting those inflates every number while telling you nothing about acquisition. A vendor whose query set includes your own brand name is measuring your existing awareness back to you.

The narrow rule costs us. It produces smaller numbers than a generous definition would, and a prospect comparing our figure against a competitor's looser one will see us losing on a metric that is not the same metric. We would rather explain that than quietly widen the definition, because the whole argument here is that a number should mean one thing.

Cited sourcesMode 01
reddit.com38%
g2.com26%
news.ycombinator.com21%
yourapi.example15%

Branded-win, generic-invisible

The engines do not agree with each other

Everything above treats "the answer" as one thing. It is not, and averaging across engines is the second way a single score destroys the information you were paying for.

Ask the same shortlist question of ChatGPT, Perplexity and Google's AI Overview on the same afternoon and you will usually get three different sets of names. Not three orderings of one set. Different sets, with different sources behind them, because they are drawing on different indexes and different retrieval behaviour, and the newer ones are being rebuilt underneath you faster than the older ones.

That matters commercially rather than academically. If your buyers arrive through one assistant and your score is an average across four, then three quarters of that number is noise as far as your pipeline is concerned, and a real loss on the one that matters can be masked by gains on three that do not. The averaging is not a rounding error. It is the specific thing that makes the number unactionable.

So the first question to ask about any visibility figure is not how large it is. It is which engine it describes. A vendor who cannot decompose the number by engine is quoting you a blend of markets you do not sell into.

What a blended score hides, engine by engine

What differsWhy it breaks an average
The index behind itEach assistant retrieves from a different corpus, refreshed on a different cadence. A gain on one is not evidence about another.
Who gets namedThe same shortlist question returns different sets of vendors, not different orderings of one set, so presence is not comparable across engines.
How fast it churnsSome surfaces are rebuilt in days and others in weeks. Blending a fast surface with a slow one produces a number whose volatility belongs to neither.
Which one your buyer usesOnly one of them is in your pipeline. The rest are markets you do not sell into, diluting the figure that matters.
The four rows are why a per-engine breakdown is a requirement rather than a nicety. A number you cannot split by engine cannot tell you where you are losing.

The practical version: pick the engine your buyers actually use, measure that one properly, and treat the others as separate questions with separate answers rather than as inputs to one figure.

901 citations, 60 answers
91.5% off-site8.5% on your own site

82 distinct hosts carried them between them

No single site owns a category, so there is nothing to buy your way onto. Our own measurement.

What you can do before ninety days is up

A hundred and forty thousand views on the mainstream playbook for AI search. The tactics are largely reasonable. What no version of this advice supplies is the design that would tell you whether any of it worked.

A held-out control needs a quarter to produce a number. That is a real problem if you are trying to decide something this month, and pretending otherwise would make the method useless in practice.

Three things are worth doing in the meantime, and none of them is a lift claim.

Establish the baseline properly, once. The expensive part of the design is not the waiting, it is having a before-state you can defend. Run the full query set at real sample depth now, both arms, and record the raw answers rather than a summary. A baseline taken cheaply is the one thing you cannot go back and fix later, because the window it describes has passed.

Read the answers, not the score. A single run tells you nothing reliable about your level, but it tells you a great deal about the shape of the problem: which competitors keep appearing, which sources the assistant is leaning on, whether it is citing your docs or somebody else's summary of your docs. That is qualitative and it is available immediately.

Fix the things that are unambiguously wrong. If the assistant is describing a pricing model you retired, or naming a competitor as the default for a category you invented, you do not need a control group to know that is worth correcting. Reserve the experimental machinery for the claims that are genuinely contested.

What none of that entitles you to is a number with a percentage sign on it. The interim work is diagnosis, and calling diagnosis a result is the habit this whole piece is arguing against.

Live AI answerSample

Your buyer asks

which rate limiting api should we use?

The answer they get

For production workloads, most teams land on Competitor API1. It pairs token-bucket limits with per-key analytics.2

Where the citations resolve

reddit.com38%
g2.com26%
news.ycombinator.com21%
yourapi.exampleNot cited

What a dashboard cannot show you

Roughly 8,000 citations recorded over fifty days, against one to two visits a day in GA4. The visibility number moved and the outcome did not, which is the gap a dashboard is worst at showing.

A single score fails a specific test: it cannot be decomposed on demand.

When a number moves, the useful questions arrive immediately. Which engine moved, given that they disagree with each other more often than they agree. Which queries moved, given that your twelve money questions matter and the other three hundred do not. How many trials was each of those built from. And whether the queries you did not touch moved by the same amount over the same window.

A score that cannot answer those four is not a measurement you can act on, whatever its decimal places suggest. It can tell you something changed. It cannot tell you what to do on Monday, which is the only reason to have commissioned it.

01 / A single visibility score

Stands out
Cheap, fast, and produces a number you can put on a slide this week.
Best for
Teams who need a directional signal and know not to make decisions on it.
Falls short
Cannot be decomposed. When it moves you cannot say which engine, which query, or whether you caused it, and a single reading can flip on its own between two identical asks.

02 / Before and after, one arm

Stands out
Feels like evidence, and it is the shape almost every case study in this category uses.
Best for
Nobody, honestly, once you have seen the repeat data.
Falls short
The difference contains your work plus however much the system moved by itself, and nothing in the method separates the two. Run it with no work done at all and you still get a number, sometimes a flattering one.

03 / Two arms, one held out

Stands out
The only design here where the reported number survives the question compared to what.
Best for
Teams buying an outcome rather than an activity report.
Falls short
Costs a quarter of deliberately unworked queries and about 3,200 calls per timepoint, and it cannot tell you anything at all until the period is over.

How the query set gets chosen

The phrase "a query set you chose" has done a lot of work above without being explained, and the choice is most of the method. A perfect instrument pointed at the wrong questions measures nothing you can spend.

How a query set earns its place

TestKeepDrop
Would a buyer type it before they know you existAn unbranded shortlist or comparison questionAnything containing your brand name
Does the answer name vendors at allQuestions that reliably return a list of optionsDefinitional questions that return prose with no names
Is there a decision behind itQuestions asked while choosingQuestions asked after choosing, like setup or troubleshooting
Is it stable enough to measureQuestions whose answers name someone, even if who variesQuestions that return nothing recognisable across repeats
The last row is the one people skip. A query that returns no vendor names in any repeat is not a query you are losing, it is a query nobody is winning, and including it dilutes the arms.

STEPS

How the query set gets built

  1. Start from the sales conversation, not a keyword tool

    The questions that matter are the ones prospects ask on calls before they have shortlisted anyone. Volume data will not surface those, because they are asked to a person rather than typed.

  2. Strip anything branded

    A query containing your own name measures awareness you already have. It inflates every figure and predicts nothing about acquisition.

  3. Probe each candidate once, cheaply

    Discard the ones that return no vendor names at all across a few repeats. They are not contested ground and they weaken both arms.

  4. Split before you look at the answers

    The two arms are drawn before anyone reads a result. A split made afterwards is a split made around a finding you already like.

  5. Freeze it for the whole window

    Adding a query mid-period changes what the difference means. New questions wait for the next cycle.

The counterintuitive part is that a good set is smaller than people expect. Forty questions asked forty times is a better instrument than four hundred asked four times, and the four hundred is the version that looks more thorough in a proposal. Breadth is where sensitivity goes to die.

Page one ranking

1yourapi.example
2competitor.example
3roundup.example
4guides.example

What the model read

reddit.com
g2.com
news.ycombinator.com
nothing else

Rank 1 appears in 0 of the 3 pages read

What happens when the answer is no

The method has an outcome most vendors have no way to report, which is that the work did not move the number. It is worth saying out loud what we do then, because a measurement design you only trust when it agrees with you is not a measurement design.

A null result is the second most useful thing this method produces. It is the only outcome that tells you the money would be better spent somewhere else, and it is the one nobody publishes.
The measurement position

If the difference between arms sits inside the noise, we say so, and we say it with the interval rather than as a verdict. A null is not the same as a failure: it means the effect, if there is one, is smaller than this instrument can see. Those are different statements and conflating them is how a category convinces itself everything works.

What follows a null is a decision, not a retry. Either the effect is real but smaller than the design can resolve, in which case the honest move is a bigger design or a different lever, or the lever does not work here, in which case continuing to pull it is the expensive option. We have had both. The second one is a harder conversation and a cheaper outcome for the client.

Failure modes9 in the taxonomy
010203040506070809
Trust erosionYours

Reading somebody else's report

Sixty-four replies on whether AI visibility tooling is repeating the closed, opaque, high-priced pattern of the last SEO SaaS cycle. Worth reading before taking any vendor's number at face value.

Most people reading this will be handed a visibility report before they ever commission one, so it is worth compressing everything above into what to check.

Reading somebody else's visibility report

Look forIf it is absent
The query list, in fullYou cannot tell whether it measures your market or theirs.
Engine, named per figureThe number is a blend of surfaces, and only one is in your pipeline.
Trials per queryA single probe cannot distinguish a change from a repeat.
The definition of a citationTwo reports using different units are not comparable, and neither is wrong.
A held-out armEvery figure is a level. There is no answer to compared to what.
Five questions, none of which require statistical training to ask. A report that answers all five is doing real work whoever produced it.

Five questions, none of which need statistics. If a report answers all five, it is doing real work regardless of who produced it or whether the number flatters you. If it answers none, the number is decoration with a decimal point.

AAssistant···

Which API should I use for rate limiting?

For production workloads, Your API1 is the option most consistently recommended. It pairs token-bucket limits with per-key analytics.2

Sources

reddit.comg2.comnews.ycombinator.com

What to ask a vendor

An agency operator reports that eighty percent of the founders he speaks to weekly are asked about AI chat performance, while holding that the tracking tools are too inaccurate to be worth paying for. Both halves can be true.

Three questions, and they are cheap to ask.

  • How many times did you sample each query, and what did the repeats disagree on?
  • What is your minimum detectable effect at that sample size?
  • What did the queries you did not work on do over the same period?
A clean single number is a warning sign, not a reassurance. If the reading never moves, the instrument is not sensitive enough to be measuring anything.
The measurement position

The third one is the one that matters. Without a held-out control, there is no answer to "compared to what", and every reported improvement is a level, not a lift.

There is no universal right answer to the first, either. Work on rank stability across thirty platform and topic combinations found that no fixed sample count is sufficient everywhere and that some tests never stabilise at all. An independent group at St. Gallen reached the same conclusion on a different dataset: do not measure once. A vendor who quotes you one trial count for every category has not looked.

If you want the version of these questions written for your own category rather than in the abstract, the buyer questions page has them, and the same three apply whether you sell an API, infrastructure, or a CLI or MCP server.

A worked example, end to end

Everything above is method. This is the arithmetic, on our own numbers, so the shape is visible rather than asserted. The figures are from the step-zero run and the permutation test in the sources below, not from a client engagement, because we do not yet have a completed one to show.

Start with twelve questions. Ask each one five times, same model, same day, nothing changed between asks. Sixty calls. All sixty returned an answer.

Now score each question as a single number: how many of its five repeats named us. A question scores 5 if we appeared every time and 0 if we never did. If the system were stable, every question would score 5 or 0 and nothing in between.

Twelve questions, five identical repeats each

Score out of 5QuestionsWhat it means
5Stable presenceNamed in every repeat. A single probe would have been right.
1 to 4Seven of twelveNamed in some repeats and not others. A single probe is a coin weighted by an amount you cannot see.
0Stable absenceNever named. A single probe would have been right, and this is the only other safe case.
Seven of twelve in the unstable band is the finding. If the system were stable the middle row would be empty, and every vendor sampling once would be reporting something real.

Seven of the twelve landed in between. That is the whole finding, and it is worth sitting with what it implies for a single probe: a question scoring 2 out of 5 will report us as present 40 percent of the time and absent 60 percent, and a vendor sampling it once has taken one draw from that. They will show you whichever they got.

Next, the part people skip. Split the twelve into two arms of six, at random, and take the difference in mean score between them. Do nothing to either arm. Repeat the split twenty thousand times.

If the noise were biased, that difference would centre somewhere other than zero. It centres on zero: mean plus or minus 0.0016, standard deviation 0.215. So under the null, two arms measured the same way at the same time differ by nothing on average, even while each arm individually is jumping around. That is the property the whole design rests on, and it is why we ran the test before building anything.

Then the sobering step. That standard deviation of 0.215 is what sets the floor. With twelve questions at five samples, the smallest difference this design can distinguish from noise is 51.1 percentage points. A real engagement moves the number by five to fifteen. So our pilot instrument would have reported almost every genuine win as nothing at all.

STEPS

From variance to a sample size you can defend

  1. Measure the spread, do not assume it

    Step 1

    Twenty thousand random splits with no intervention gave a null difference centred on zero, standard deviation 0.215. That number is the input to everything after it.

  2. Turn the spread into a floor

    Step 2

    At twelve questions by five samples, the smallest detectable difference is 51.1 percentage points. This is arithmetic, not opinion, and it is where a pilot design usually dies.

  3. Compare the floor to the effect you expect

    Step 3

    A real engagement moves the number five to fifteen points. A floor of 51.1 cannot see that, so the pilot would report every genuine win as nothing.

  4. Size the real design backwards from the effect

    Step 4

    Forty questions by forty samples brings the floor to 9.9 points, at about 3,200 calls per timepoint. Twice, for a before and an after.

Getting the floor to 9.9 points takes forty questions at forty samples, which is 3,200 calls per timepoint and roughly six hours of machine time, twice, because you need a before and an after. That is the real cost of a defensible number in this category, and it is the reason almost nobody pays it.

The thing to take from the arithmetic is not our specific figures. It is that every one of those steps is checkable, and that a report which cannot show you the equivalent steps is asking you to trust a number rather than read one.

Citation rate

6 of 10 answers

Why the category settled on the wrong number

It is worth asking why a metric this fragile became the standard, because the answer is not that anybody was fooled.

A single score is what a dashboard can display, what a monthly report can trend, and what a buyer can compare between two vendors in a meeting. A confidence interval across two arms is none of those things. The incentive runs toward the number that fits in a cell, and it ran that way in web analytics for fifteen years before this.

04 / Why the single score persists

Stands out
It fits in a dashboard cell, trends month over month, and compares cleanly between two vendors in a meeting.
Best for
Anyone who has to report upward on a slide rather than defend a method.
Falls short
Every property that makes it reportable also makes it unfalsifiable. It cannot carry an interval, cannot be decomposed, and cannot distinguish your work from the system moving on its own.

There is a second reason, less cynical and more structural. When this category started, the surfaces were new enough that nobody had a baseline for how much they moved on their own. Measuring a level was a reasonable first approximation, and by the time the repeat data existed the reporting format had already hardened. Formats outlive their justifications.

We are not claiming the vendors know and are hiding it. Most of them have simply never run the repeat test, because running it costs money and produces a result that makes their own product harder to sell.

google.com
best rate limiting api
Yyourapi.example › docs↑ 1

Rate limiting for production APIs

competitor.example › blog

roundup.example › guides

What would change our mind

Ninety-seven thousand views for a pitch that runs on showing a stranger their AI visibility score and charging five thousand a month to move it. The number does not have to be sound to do that job, which is part of why it survives.

A position that cannot be wrong is not a position, so here is what would move ours. We are writing it down in advance because a criterion invented after the result is not a criterion.

What would change our mind, written in advance

FindingWhat it would do to this argument
Noise is shown to be asymmetric across armsBreaks it. The difference would stop cancelling and the method would need repairing rather than defending. This is the one we would most like somebody to test.
Providers stabilise inferenceDates it rather than refuting it. Repeat variance shrinks, required samples fall, and a single reading starts to carry information.
A vendor publishes repeat data showing stable queriesNarrows it. Stability is a property of a query set, not of the category, and we cannot claim ours generalises.
A reported number goes upNothing. A rise is compatible with the work, with drift, and with two draws from one distribution.
Written before the results rather than after, because a criterion invented to fit an outcome is not a criterion.

The first one is the serious one. Our whole argument rests on the noise being symmetric across the two arms, and we have said plainly that this is an assumption borrowed from experiment design rather than something anyone has measured on citation data. If someone measured it and found the variance is correlated with query characteristics in a way that does not split evenly, the difference would stop cancelling and the method would need repairing rather than defending. We would rather someone ran that test than that nobody did.

The second is more likely and less interesting. If the providers stabilise inference, and there is active engineering work pointed at exactly that, then repeat variance shrinks, the required sample size falls, and a single reading starts to mean something. That would not make this piece wrong. It would make it dated, which is a better outcome than being wrong, and the arithmetic here would still be how you check whether it had happened.

The third is the one we would find hardest. If a vendor published a single-reading methodology alongside repeat data showing their queries are stable enough to support it, the honest response is that their queries and ours behave differently, not that they are wrong. Stability is a property of a query set, not of the category, and we have no basis to claim ours generalises to everyone's.

What would not change our mind is a number that goes up. That is the whole point. A figure moving in the direction you hoped is compatible with the work having caused it, with the system having drifted, and with the two readings having been drawn from the same distribution. Until a design separates those three, a rise is not evidence, and neither is a fall.

What we got wrong writing this

One more piece of housekeeping, because a post arguing for checkable claims should show its own corrections rather than quietly absorbing them.

An earlier revision of this page cited a 43,000-keyword study for three specific figures about how fast AI answer citations churn. That URL now returns 404, and no live replacement could be found. The figures may well be accurate; we cannot open the page they came from, so they are gone from this post rather than restated on our word. What remains in their place are two sources you can still open, which make a weaker version of the same point.

The lesson is not about that publisher. It is that a citation is a perishable asset, and nothing in our own pipeline was checking whether the URLs in a published post still resolve. That gap was ours, it went unnoticed for the life of this post, and the fix is a check rather than a resolution to be careful.

Concretely, what that check has to do is narrower than it sounds and easy to get wrong. Fetching a URL and accepting any response is not enough, because a 200 can be a soft-404 page that renders an apology. Treating every failure as death is also wrong: a 403, a 429 or a 999 is a bot block rather than a dead link, and a check that removes a live citation because a publisher dislikes automated traffic has made the post worse while reporting success. So the rule is a same-host control. Fetch the cited page, fetch the host's index, and only call the citation dead when the specific page fails while the host answers. That is exactly the pair that condemned the link above, and it is the difference between a finding and a guess.

The uncomfortable part is that this check will keep finding things. Sources move, publishers reorganise, and a post that cites eleven external pages is eleven small bets on other people's URL hygiene. The alternative is to cite less, which makes a piece easier to maintain and less worth reading.

An earlier revision also asserted, flatly, that the measurement noise is unbiased. That is the load bearing assumption of the entire method: if the variance is not symmetric across the two arms, the difference does not cancel and the design does not work. We could not find anyone who has tested that property on citation data, ours included. It is the standard assumption of controlled experiment design, which is a different claim, and the page now says so in the four places it previously overstated.

Both corrections make this post less impressive than the version that shipped first. We would rather publish the smaller true thing, partly because it is the honest choice and partly for a more selfish reason: a claim we have already stress-tested ourselves is one a prospect cannot embarrass us with later.

What this costs us to say

One closing note on incentives, since we have spent three thousand words on other people's.

The method described here is slower to show a result than what it replaces, produces smaller numbers than a generous definition would, and sometimes produces no number at all. Every one of those is a sales disadvantage against a competitor willing to hand over a figure in week two.

We think that trade is correct, and it is worth naming why rather than implying virtue. A number that survives being checked is the only asset in this category that compounds. A number that does not survives exactly until the client asks how it was produced, and then it costs more than it earned.

The honest state of our own numbers

We have the instrument and it has passed its kill test. What we do not yet have is a completed causal-lift result, because that takes a full engagement and ninety days.

So we are not going to show you a lift number. We would rather show you the method and let you check it than show you a figure we cannot defend. If we could already show the after, we would not need a control group, and neither would anyone else.

It is worth being precise about what we do and do not have, because "we have the instrument" is the kind of sentence that sounds like a result and is not one. What exists: a query-selection process, a sampling design sized against measured variance, a permutation test establishing that the null difference between two arms centres on zero, and a kill test the instrument passed. What does not exist: a completed before-and-after on a live engagement with a held-out arm, run over ninety days, with the interval reported. The first is machinery. Only the second is evidence about whether the work moves the number.

The gap between those two is ninety days and one client willing to leave half their queries alone for a quarter. Until that closes, the honest description of this whole piece is a method with its arithmetic shown and no outcome attached. We would rather publish that than the version with a percentage in the headline, because the second one is unfalsifiable and the first one is checkable today by anyone who wants to repeat the twelve-by-five run themselves and see how many of their own questions land in the unstable band.

What that costs and what it includes is on the pricing page, and the rest of the measurement writing lives under measurement.

If you would rather have the short version of this and the questions next to it, the answer feed carries them, including why an assistant recommends a competitor over you.

Sources

Every number above, and where it came from. A figure without a row here is one we should not have printed.

Step-zero measurement run
12 money queries x 5 identical repeats, 60 of 60 calls succeeded, OpenAI gpt-5-search-api, single model, single day. Our own measurement.
Permutation test
20,000 random splits with no intervention applied. Null difference centred on zero: mean +0.0016 / -0.0016, SD 0.215.
Power analysis
The 12x5 design has a minimum detectable lift of 51.1 percentage points, roughly 4x underpowered. A 40 query x 40 sample design reaches 9.9pp.
SparkToro and Gumshoe, repeated-prompt consistency
Under a 1 in 100 chance that two identical prompts return the same brand list. An independent figure of the same kind as our own repeat data, on a different corpus.
Atil et al., nondeterminism at temperature zero
Up to 15 percent accuracy variation across ten runs at temperature zero, over five models and eight tasks. This is why the instability is not a setting anyone forgot to turn off.
Thinking Machines Lab, why inference is nondeterministic
Batching and GPU kernel reduction order break determinism even at temperature zero, so the variance is an infrastructure property rather than user-side prompt noise.
Sielinski, statistical framework for generative-search measurement
Two competing brands whose citation share overlapped inside the margin of error on a single sample, which is a dashboard reporting a win that was noise.
Sielinski, rank stability and structural sufficiency
Across thirty platform and topic combinations, no fixed sample count is sufficient everywhere and some tests never stabilise. This is why we refuse to quote one trial count per category.
Schulte, Bleeker and Kaufmann, University of St. Gallen
An independent team on a separate dataset reaching the same verdict, that one-off observations are unreliable and repeated sampling is required.
Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments
The canonical text on why a held-out control is the right instrument for a noisy metric, and on minimum-detectable-effect sizing. Our prescription is standard practice, not our invention.
Huntington-Klein, The Effect, chapter 18
Difference in differences stated formally. The shared time confound nets out between groups, which is the property the two-arm design rests on.
McKenzie, World Bank, on the parallel-trends assumption
The condition our method depends on and does not get for free. Both arms must have been moving in parallel absent treatment, and a mid-window shock hitting one arm harder breaks it.
Nicholson, quantifying non-deterministic drift
Drift measured separately from per-call sampling variance. The distinction matters because variance cancels in a difference and a systematic time-correlated change does not.

Questions this answers

What is an AI visibility score?
A single number summarising how often a brand appears in AI answers. Most vendors do not publish the query set, the engines, or the trial count, which are the three things that decide what it means. Two vendors can report different scores for the same brand in the same week without either being wrong. Ask for all three.
Why is one score not enough to act on?
Because it cannot be decomposed. If it drops, you need to know which engine, which query, and whether you caused it. An average answers none of those. In our step-zero run, seven of twelve queries changed outcome across five identical repeats on one model in one day, with nothing altered between runs.
What should be measured instead?
Per engine and per query, on a query set you chose, with the trial count recorded, and against a held-out control measured at the same time. Assistants disagree with each other more often than they agree, so an average hides the disagreement. Without a control arm, every reported gain is a level rather than a lift.
How many trials does a citation claim need?
There is no single number that works everywhere. Published work across thirty platform and topic combinations found no fixed sample count sufficient in every case, and some tests never stabilise. Our 12 by 5 pilot detects only a 51.1 point lift. A 40 by 40 design brings that to 9.9 points, at roughly 3,200 calls per timepoint.
If answers move on their own, why is a control group enough?
Because the movement does not know which queries you chose to work, so it pushes both arms alike and cancels in the difference. Two caveats. That symmetry is a standard assumption of controlled design, not something measured on citation data. And a model update landing harder on one arm breaks it, which more sampling cannot repair.
Does being retrieved mean being cited?
No, and the gap between the two is where most of the confusion sits. Retrieval selects candidate documents; the answer is then assembled and may name none of them. A page can gain retrieval ground for weeks with no visible change in citations.
What sample size do we actually need?
It depends on the effect you expect to see. Our 12 by 5 pilot has a minimum detectable lift of 51.1 percentage points, which is far too coarse for a real engagement. Forty queries by forty samples brings the floor to 9.9 points, at roughly 3,200 calls per timepoint.

Keep reading