New: the nine ways developer tools go invisible in AI answers.
citon
AI citations

AI Visibility Tools Show a Score, Not Whether It Was You

Ten guides rank the best AI visibility trackers. None of them tells you whether the score moving was your AEO spend or the model changing its mind on its own.

Citon23 min read

AI visibility tools show you a score. They can't tell you if it was you. A decision framework for buyers who already own a monitoring tool and still cannot answer whether their AEO spend worked.

Short answer

Can an AI visibility tool prove your AEO spend worked?

No. An AI visibility tool reports a score, a count of how often a brand appears in AI answers this week versus last week. It cannot tell you whether that movement was caused by work done on your behalf or by the underlying model changing its own answers, because a single reading has no comparison point. Proving causation needs a second thing the dashboards do not ship: a held-out control, a set of identical queries deliberately left unworked across the same window, so whatever the system did on its own can be subtracted out of whatever the score reports. Without that second arm, a rising number and a falling one are both compatible with doing nothing at all.

Ten guides currently rank the best AI visibility tools. Every one of them compares dashboards on features, coverage, and price, and every one of them assumes the buyer's open question is which tool to buy. It is not, not anymore. The buyers showing up in r/SEO and r/DigitalMarketing already own a monitoring tool. Their question is what the number it reports actually proves, and the honest answer is less than the dashboard implies.

What an AI visibility tool actually measures

An AI visibility tool asks a fixed or semi-fixed set of queries against one or more AI systems on a schedule, counts how often a target brand is named in the responses, and reports that count as a score, sometimes rolled into a single index, sometimes broken out by engine. That is a real measurement of appearance frequency. It is not a measurement of cause. Two readings taken a month apart tell you the count went up or down; they say nothing at all about why, because a single trend line has no comparison point built into it.

At a glance

What a score answers, and what it cannot

Question a buyer actually hasCan a visibility score answer it
Are we cited more than we were last quarterYes. That is the one thing the number is built to report.
Did our AEO spend cause that changeNo. A single reading has nothing to compare itself against.
Would the number have moved without usNo. That requires a second, unworked arm the dashboards do not run.
Which query, which engine, moved the totalSometimes, if the tool decomposes by query. Most ranked-listicle tools lead with the rolled-up number instead.
Is the reading stable if we ask again tomorrowRarely tested. Our own repeat run found 7 of 12 identical questions changed answer between identical asks.

The gap matters more than it sounds like it should, because the buyer's actual question is almost never "what is our score." It is "did the money we spent cause this," and a rolled-up count answers a question adjacent to that one rather than the one being asked. A score can rise because an agency did real work. It can also rise because the underlying model changed how it answers a class of query, independent of anything any vendor did. From the outside, both cases produce the identical chart.

A score that never moves is not a calm result, it is an instrument too coarse to be measuring anything. A score that moves proves even less, unless you know what it would have done on its own.
The measurement position

The category's own vendors describe the product this way when they are being precise about it. A representative comparison of monitoring tools, Backlinko's own rundown of LLM tracking products, frames every entry on the list as answering "how often are we mentioned," never "why did the mention count change." That framing is not a marketing shortcoming unique to one vendor. It is close to a structural limit of what a monitoring product can promise, because answering the causation question requires running a second, deliberately unworked arm of the same query set, and a monitoring subscription has no natural place to put one. A dashboard that tracked a control group would need a customer willing to pay to leave part of their own coverage untracked on purpose, which is a harder thing to sell than a rising number.

A second, quieter problem sits underneath the causation gap: most visibility tools do not publish how the underlying query set was chosen, how many times each query was actually asked per reporting period, or how the vendor decided which mentions count as a citation versus an incidental name-drop. Those choices change the reported score independently of anything a buyer did. Two vendors running the identical brand through the identical models can report meaningfully different numbers purely from differences in query curation and counting rules, before either one gets anywhere near the causation question this post is about.

Why are buyers asking this question right now?

The evidence that this has become a live buyer concern, rather than a hypothetical one, is not hard to find. A buyer on r/SEO described paying an AEO agency a four-figure monthly retainer for a local service business and watching the citation count move without any way to tell whether the movement reflected the agency's work or the model doing what it would have done anyway. A hundred and three comments followed, the overwhelming majority some version of the same unanswered question.

A buyer describes a four-figure monthly AEO retainer and no way to separate the resulting citation count from what would have happened with the account left alone. 103 comments, most of them the same question restated.

The pattern repeats across a wider set of threads than one frustrated buyer. An AEO agency operator ran their own AMA describing a flat monthly retainer and a reported visibility number; the comments underneath, rather than the AMA's own answers, are where the real question surfaces, over and over, in slightly different words: how do you know the number would not have moved on its own. Nobody in the thread gets a satisfying answer, including the operator running the AMA.

An AEO agency operator's own AMA. Read the comments rather than the answers, the recurring ask is proof the reported number is attributable, and the AMA does not supply it.

A third thread takes the same question to a wider, less agency-adjacent audience. A buyer whose own marketing agency is upselling AEO on top of an existing SEO retainer asks a general marketing subreddit whether the additional spend is actually worth it. A hundred and thirty-one comments split, roughly, between people who think the category is real and people who think the measurement underneath every AEO pitch is too thin to trust. Both groups are responding to the same underlying absence: nobody in the thread can point to a vendor who has shown their own reported number survives a real control.

A buyer takes their own agency's AEO upsell to a wider audience. 131 comments, split between "worth it" and "how would you even know."

The uncertainty is not confined to buyers on the receiving end of an AEO pitch. Practitioners on the selling side ask a version of the same question among themselves, in a thread asking whether GEO work is worth doing before a client has requested it at all. The thread is not cynical about the category, most participants think generative-engine visibility is a real and growing concern. What it lacks, from people whose job is to sell this exact service, is any agreed answer to how they would show a client the work caused anything. If the people selling the measurement cannot describe how they would prove it, a buyer evaluating a finished pitch deck is unlikely to get a clearer answer than the sellers have themselves.

Practitioners, not just buyers, debating whether GEO spend is defensible before a client has asked for it. Even people selling the work are unsure what evidence would settle the question.

The three questions a score cannot answer

Strip the marketing language off any AI visibility dashboard and it is answering one question well and two questions not at all. It answers "how often are we named" accurately, assuming the query set and the sampling are honest. It does not answer whether that count would have been different with nobody working the account, and it does not answer how much confidence the number deserves given how few times each query was actually run.

Three questions a score cannot answer, and what a real answer requires

QuestionWhat a score-only vendor saysWhat a real answer requires
Did our spend cause the changeShows you the changeA held-out arm run the same window, so the system's own movement can be subtracted
Would it have moved without usDoes not askThe same instrument run with nobody working the account, for comparison
How confident should we be in the numberReports a single readingA trial count large enough that the reported difference is bigger than the noise floor
A vendor that answers the first two questions with a shrug has, by construction, never separated their work from the system's own drift.

None of this makes a visibility score worthless. It makes it a narrower instrument than the pitch decks built around it suggest. A speedometer tells you how fast the car is going; it does not tell you whether the road is sloped downhill. Both facts matter to a driver, and only one of them is on the dashboard. The AEO category's dashboards are, almost without exception, built to answer the speedometer question and silent on the slope.

The confidence question deserves its own attention, because it is the easiest of the three to fix and the one most consistently skipped anyway. Trial count, how many times a given query was actually asked before the reported figure was computed, is not a subtle statistical nuance. It is closer to a label on the side of the box: a reading based on one ask per query carries almost no information about what the underlying system does on average, while a reading based on five or ten asks per query starts to. A dashboard that reports a score with no stated trial count is not necessarily hiding a small one, but a buyer has no way to tell the difference, and the category's convention is not to state it either way.

Pull the current search results for the term "ai visibility tool" and the pattern is stark rather than subtle. Ten of ten results are ranked buying guides or vendor product pages, comparing monitoring dashboards on feature lists, pricing tiers, and engine coverage. Not one of the ten addresses the question a buyer asks after they already own a monitoring tool, which is whether the number it reports is attributable to anything at all.

What the current top-10 results for "ai visibility tool" actually are

ResultShapeAddresses causation
Frase, 10 Best AI Visibility ToolsRanked listicleNo
Zapier, 8 best AI visibility toolsRanked listicleNo
Profound, 18 Best AI visibility tools for agenciesRanked listicleNo
Semrush, AI Visibility overviewProduct pageNo
Backlinko, 5 AI Visibility Tools to Track Your BrandRanked listicleNo
Ten of ten results in the live search are pre-purchase buying guides for a monitoring dashboard. None targets the buyer who already owns one and wants to know whether the number moving was their AEO spend.
Ten ranking articles for "ai visibility tool," all comparing monitoring dashboards, none addressing the post-purchase causation question.
The current search results for "ai visibility tool." Ten of ten are pre-purchase buying guides. Zero address what a buyer asks after they already own one.

That gap is not an oversight in any single article. It is a structural feature of the category these guides are writing for. A monitoring-tool comparison exists to help a buyer choose between dashboards, and every dashboard on the list reports the same kind of score, a count with no control arm. Ranking the dashboards against each other on features cannot surface a problem that every dashboard on the list shares. The gap only becomes visible once a buyer stops asking which tool to buy and starts asking what the tool's own number actually proves, which is exactly the question none of the ten articles is written to answer.

A buyer doing real homework on the category already gets this far without any outside prompting. One buyer described spending six months testing four named monitoring vendors end to end before writing up the results publicly. That is the comparison work the ranked listicles claim to do on a buyer's behalf, done by hand, by someone who still landed on the scoring problem rather than a features problem once the testing was finished.

There is also a plainer, less charitable explanation worth naming, because it is the one a buyer is least likely to hear from a vendor directly. A monitoring dashboard is genuinely easier to sell than a held-out control. A rising number fits in a single screenshot, updates on a schedule a customer can check without help, and never requires telling a paying client that a slice of their own query coverage was deliberately left untouched for a quarter. None of that makes the dashboards dishonest. It does mean the category's commercial incentives point away from the harder, slower design this post is describing, and a buyer who assumes the market will supply proof on its own, because proof would obviously be better, is assuming against those incentives rather than with them.

The domain-authority shape of the current search results reinforces the same point from a different angle. Every one of the ten ranking guides for "ai visibility tool" is published by a site with a substantial, years-old backlink profile, the kind of authority a new domain cannot out-rank on a bare "best of" list at launch regardless of how good the underlying argument is. That is not a reason to avoid the keyword. It is a reason not to compete on the ranking guides' own terms, because the fight they have already won is a different fight from the one a buyer with an unanswered causation question actually needs settled.

Is a rising score proof of your work, or the system moving on its own?

This is the question a single reading structurally cannot answer, and it is worth being precise about why, because the honest answer is not "it depends on the vendor," it is "a single number cannot answer it regardless of the vendor." Our own step-zero measurement run tested twelve money queries five times each against one model on one day, with sixty of sixty calls succeeding. Seven of the twelve questions changed their answer between identical repeats asked minutes apart, with no intervention of any kind applied between them.

That instability is not a defect in our own instrument. It shows up independently in research nobody at this company ran. SparkToro found under a one-in-a-hundred chance that two identical prompts return the same brand list. Atil et al. measured up to fifteen percent accuracy variation across ten identical runs at temperature zero, across five different models and eight different tasks. Both findings describe the same underlying property a rolled-up visibility score inherits by default, whether or not the vendor reporting it has ever measured it directly.

The mechanism behind the instability is worth naming plainly, because "the model is random" undersells what is actually happening. Batching behaviour on the provider's own infrastructure, kernel reduction order on the GPUs serving a request, and routine model updates a provider ships without a version bump all change an answer between two calls that look identical from the outside. None of those three causes has anything to do with whether a brand's AEO work is any good. A visibility score cannot distinguish "the model changed" from "the work changed," because both produce the identical artifact: a different count on the second reading than the first.

From the field

The instability behind a single score is measured independently of our own work

SparkToro found under a 1 in 100 chance that two identical prompts return the same brand list. Separately, Atil et al. measured up to 15 percent accuracy variation across ten identical runs at temperature zero, across five models and eight tasks. Neither team was testing an AEO vendor's claims; both were describing the same property a single visibility reading inherits by default.

SparkToro and Gumshoe, repeated-prompt consistency

A permutation test we ran separately makes the practical consequence concrete. Twenty thousand random splits of the same query set, with no intervention applied to either half, produced a null difference centred almost exactly on zero, with a standard deviation wide enough that a real, modest AEO effect would sit comfortably inside the range that pure noise already produces on its own. A score that moved inside that range proves nothing about whether anyone did any work at all.

Put a number on what that means for the reader deciding whether to trust their own agency's slide. If the reported change since last quarter is smaller than roughly two standard deviations of that null distribution, meaning the swing that turns up with nobody touching the account at all, it is not distinguishable from noise using a single before-and-after reading. Nothing about a bigger swing proves causation either, since a large enough system-side shift, a model update landing mid-quarter, can move a single-arm reading by more than that on its own. The only way to tell the two apart is the same held-out arm the rest of this post keeps returning to, run over the same window, on the same query set, reported as a difference rather than a level.

The buyer's actual question is never "what is our score." It is "did the money I spent cause this," and a dashboard cannot answer a question it was never built to ask.
The measurement position

The held-out control, explained without the arithmetic

The fix for this is not a better dashboard. It is a different measurement design, one that almost none of the ten ranking guides for "ai visibility tool" mention, because it is not a feature a monitoring product can ship on its own. It requires the buyer, or whoever is running the measurement, to deliberately leave part of the query set unworked for the length of the engagement.

This is not a novel idea invented for AEO. Kohavi, Tang and Xu's Trustworthy Online Controlled Experiments lays out the same logic for any noisy online metric: a held-out control is the standard instrument, and sizing the minimum detectable effect before trusting a result is standard practice, not a boutique add-on. What is unusual about the AEO category is not the method, it is how rarely a monitoring product applies it. The book's prescription predates generative search by years; AEO measurement is simply the newest metric it happens to apply to cleanly.

STEPS

The two-arm design that answers the causation question

  1. Split the query set before any work starts

    Day 0

    The queries that matter are divided into two arms before anyone reads a result, so the split cannot be drawn later around a finding someone already likes.

  2. Measure both arms at once, same instrument

    Day 0

    Same prompts, same day, same trial count, both halves. A baseline read at two different times is two different baselines, not one.

  3. Work one arm only, for the whole window

    Days 1 through 90

    The held-out half is left alone on purpose. This is the expensive part, and it is the part that makes the eventual result attributable rather than merely correlated.

  4. Re-measure both arms, identically

    Day 90

    Same instrument as day 0. Changing the prompt wording or the trial count between timepoints silently changes what the reported difference means.

  5. Report the difference between arms, not either arm's level

    Day 90

    Whatever the underlying system did on its own over the window, it did to both halves equally, so it subtracts out of the difference and only the attributable part remains.

A two-arm measurement design, one worked arm and one held-out control, reporting the difference between them rather than a single level.
The design that answers the causation question. One arm worked, one arm deliberately left alone, the same window, the difference reported instead of either arm's raw number.

The logic is simple even though the discipline to run it is not. Split the queries that matter into two groups before any work starts, so the split cannot be drawn later around a result someone already likes. Work one group. Leave the other alone, completely, for the whole window, which is the expensive and unglamorous part, because it means telling a client you deliberately did not touch a portion of their query set on purpose. Re-measure both groups at the end with the identical instrument used at the start. Report the difference between the two groups rather than either group's raw level, because whatever the underlying model did on its own over that window, it did to both groups equally, and that shared movement cancels out of the difference.

A single visibility score compared against a held-out-control proof design, showing what each can and cannot answer.
A score and a proof are different instruments answering different questions. Ten ranking guides sell the first. Almost none mention the second exists.

A single visibility score cannot do any part of this, not because the vendors selling it are being dishonest, but because the product category was built to answer a different, narrower question. A trifecta comparison makes the trade-off explicit rather than implied.

01 / A single visibility score, no control

Stands out
Cheap, fast, and produces a number a buyer can put on a dashboard the same week.
Best for
A team that wants a directional read and already knows not to make a spend decision on it alone.
Falls short
Cannot separate a buyer's own AEO work from the system's own drift. A rising number and a flat account left untouched can produce the identical reading.

02 / Before and after, one arm worked

Stands out
Feels like evidence, and it is the shape most AEO case studies in this category use.
Best for
A team that has not yet seen how much a score can move with no intervention applied at all.
Falls short
The reported difference mixes the buyer's work with however much the system moved by itself, and nothing in the method separates the two.

03 / Two arms, one held out as a control

Stands out
The only shape here where the reported number survives the question, compared to what would have happened anyway.
Best for
A buyer paying for an outcome, not an activity report.
Falls short
Costs a deliberately unworked slice of the query set for the whole window, and produces nothing you can report until that window closes.

How the query set itself decides whether the number means anything

A held-out control fixes the comparison-point problem, but it inherits a second problem the visibility-tool category is just as quiet about: which queries go into the measurement in the first place. A query set built carelessly can make a two-arm design report a confident-looking difference that still means nothing, because the queries were never contested ground to begin with.

The queries that matter are the ones a real buyer would type or ask before they have shortlisted anyone, the pre-purchase, comparison-stage questions rather than branded searches for a company that already knows it exists. A query containing a brand's own name measures awareness that brand already has; it inflates every reading and predicts nothing about whether a stranger would find that brand through an AI answer. Stripping branded queries out of the set is the first and cheapest filter, and it is also the one most likely to be skipped, because branded queries are the ones most likely to return a favourable-looking mention.

The second filter is harder to apply and easier to skip under time pressure: discarding queries that return no vendor names at all across several repeats. An uncontested query, one no AI system currently answers with any brand name, cannot move in either direction from an AEO engagement, because there is no existing citation behaviour for the work to change. Including uncontested queries in a measurement set does not bias the result toward a false positive or a false negative on its own, but it does dilute the signal, making a real effect on the contested queries harder to see once it is averaged against queries that were never going to move regardless of the work done.

The two filters together are why a measurement built for this purpose looks different from a keyword list built for search-volume research. Search-volume tools optimise for scale, more terms, wider coverage, ranked by how many people type them. A held-out-control query set optimises for the opposite: fewer terms, deliberately curated for contestedness, sized to what the trial-count arithmetic in the previous section actually needs rather than to how many keywords a research tool happened to surface.

What buyers actually say when the spend didn't show results

The buyer-side evidence for this gap is not limited to people complaining about a specific agency. It shows up in how buyers describe shopping the monitoring-tool category itself, independent of any agency relationship at all. One buyer's own account of the search, published on LinkedIn, describes months spent hunting for a single AI visibility tool that covered generative-engine tracking, citation tracking, and referral analytics in one place, and never finding one, because every candidate answered the counting question and none of them answered the attribution question the buyer actually cared about.

The broader category-level evidence points the same direction from a completely different angle. Agencies are now building AEO and GEO content as a standing line item rather than a one-off engagement, visible enough that it reached a wide general-marketing audience outside the SEO-specific communities where this argument usually stays contained. The spend is real, it is growing, and none of that growth on its own says anything about whether any specific agency's reported number is attributable to their work.

Ask any AEO vendor what their number would have read with nobody working the account. If the answer is a shrug, the number they did report was never evidence of their work in the first place.
The measurement position

Neither piece of evidence indicts AEO as a category, and it would be a mistake to read either that way. Agencies building GEO and AEO content as a standing line item is a reasonable response to a real shift in where buyers now look for answers, and a buyer shopping the monitoring-tool market for six months before committing is doing exactly the diligence a real purchase deserves. What neither piece of evidence shows is a vendor, agency, or in-house team demonstrating that their own reported number survives a real control. The gap this post is describing sits between "AEO spend is a reasonable line item" and "this specific vendor's reported score proves their spend worked," and almost every public conversation about the category conflates the two.

What "proof" would actually look like

If a rolled-up score is not proof, it is worth being specific about what would actually count as proof, rather than leaving the bar undefined. Proof, in this context, is a reported difference between a worked arm and a held-out control arm, run over the same window with the same instrument, large enough in sample size that the difference clears the noise floor a permutation test on the same query set would establish. Anything short of that is a level, not a difference, and a level cannot rule out the system having produced the identical number on its own.

What would change our mind, written before anyone reads it

FindingWhat it does to this argument
A vendor publishes a real held-out-control result showing their reported score tracks causationNarrows it. That vendor has closed the gap this post describes, and buyers should ask every other vendor for the same evidence.
Model providers stabilise inference enough that a single reading stops flipping on identical repeatsDates it rather than refutes it. A stable system needs a smaller control arm, not none at all.
A buyer's reported score rises and stays risen for a full quarter with the agency pausedNothing on its own. A single post-hoc observation is exactly the uncontrolled pattern this post argues against, whichever direction it points.
Written in advance rather than after a result, because a criterion invented to fit an outcome is not a criterion.

This bar is deliberately falsifiable, stated before any result rather than shaped around one after the fact. A vendor who publishes a real held-out-control result showing their reported score tracks causation has done the work this post describes, and buyers should hold every other vendor to the same standard rather than accepting a rolled-up number because it is the one that ships fastest.

Stating the bar this precisely also means being precise about who owns the burden of proof. It does not sit with the buyer to disprove a vendor's number; it sits with whoever is claiming causation to produce the comparison that would support it. That is a reversal of how most AEO pitches are structured today, where a rising chart is presented and the buyer is left to either accept it or argue against a number they have no comparable figure to argue against. Flipping the burden costs the vendor more work up front and gives the buyer something falsifiable to evaluate instead of something to take on faith.

When a visibility score is still useful

None of this is an argument that a visibility score is worthless, and it would be dishonest to write it as one. A single reading is a reasonable directional signal for a team that already knows not to make a spend decision on it alone, the same way a single poll result is a reasonable directional signal for an election nobody is deciding based on that poll in isolation. The failure mode is not using the number. It is trusting it to answer a causation question it was never designed to answer, and then renewing or cancelling a retainer on that basis.

A buyer who already owns a monitoring tool does not need to replace it to close this gap. The dashboard keeps doing the job it is good at, counting appearance frequency on a schedule. What closes the gap is adding the second arm alongside it, not instead of it, which is a measurement decision rather than a purchasing decision.

There is a useful test for whether a given use of a visibility score has crossed from "directional signal" into "unearned proof." Ask what decision the number is being used to justify. A weekly check-in that answers "are we still showing up at all" is well within what a single reading can support. A renewal decision, a budget increase, or a case study claiming a specific percentage lift is asking the number to do work it was never built to do, and that is exactly the point at which a second, held-out arm stops being optional rigor and becomes the minimum bar for the claim being made.

The same reasoning applies to a buyer deciding whether to renew an existing retainer versus start fresh with a new vendor. A renewal decision made off a single rising number is exposed to the identical blind spot a first purchase is: neither the buyer nor, in most cases, the vendor actually knows whether the number would have risen anyway. The right question at renewal time is not "did the score go up," it is "has anyone run the comparison that would tell us whether we caused it," and for most AEO engagements today the honest answer is no, because the comparison was never designed into the engagement from the start.

What should you ask your AEO vendor to prove it worked?

A buyer does not need to run the statistics themselves to raise the standard of what they accept as evidence. Three questions, asked directly, separate a vendor who has actually measured their own causal effect from one who has only ever reported a level. What would the number have read with nobody working the account. How many times was each query repeated, and on how many distinct queries. Is the figure being reported a raw level or a difference against a control that was run the same window.

These three questions are deliberately ordered from hardest to answer to easiest. A vendor who has genuinely run a held-out control can answer the first question with a real number, the control arm's own reading, rather than a guess. A vendor who has not run one can still answer the second question honestly, since trial count is a fact about their process regardless of whether they used it correctly, and a vendor unwilling to answer even the third question, whether the figure is a level or a difference, is declining to state what kind of number they are actually showing a client. That refusal is itself informative, and a buyer is entitled to treat it as one.

A vendor with a real answer to all three has done the work this post describes, whether or not they call it a held-out control by name. A vendor who answers with a shrug, or with a single impressive-looking chart and no mention of a control arm, has reported a score, not proof, and the distinction is the entire argument of this post. The chart may still be true. It is simply not evidence of what it is being presented as evidence of.

The gap between a score and a proof is not a reason to distrust every AEO vendor, and it is not a reason to stop tracking answer engine optimization performance. It is a reason to ask a sharper question before renewing a retainer built on a number nobody has separated from the system's own drift. Our own work on measuring answer engine optimization lift starts from the same instrument described in this post: a held-out arm, run alongside whatever monitoring tool a buyer already owns, reporting a difference instead of a level. The same discipline sits behind everything we publish on our measurement research, including the step-zero run and permutation test cited throughout this post, first published in why a single AI visibility score is noise. It is also the same lens we bring to our wider AI citations research more broadly, and to the questions developer tool buyers ask AI before they commit budget to a monitoring dashboard, an AEO retainer, or both. If the open question for you is which of the two to buy first, do you need an AEO agency or just an AI visibility tool answers it directly. Current pricing for that work is scoped per engagement rather than sold as a flat retainer, for the same reason a held-out control cannot be sized without first seeing the query set it will run against.

None of this requires taking our word for any of it. The step-zero run, the permutation test, and the arithmetic behind the trial-count floor are all published with their raw numbers rather than summarized as a single reassuring claim, on the same domain this post lives on. A buyer who wants to check whether the argument holds together can read the underlying measurement directly, rather than trusting a blog post's characterization of it, which is the same standard this post has been asking AEO vendors to meet throughout.

Sources

Every number above, and where it came from. A figure without a row here is one we should not have printed.

r/SEO, "spent a four-figure sum on an AEO agency for a pool business and got nothing"
58 upvotes, 103 comments, 2026-07-02. A buyer describes paying an AEO agency a four-figure monthly sum with no way to tell whether the resulting citation count would have happened anyway.
r/aeo, AMA with an AEO agency operator on a flat monthly retainer
87 upvotes, 66 comments, 2026-06-12. The operator's own AMA thread, read for what the buyers in the comments ask that the AMA itself does not answer, namely how the reported number is separated from the system's own drift.
r/DigitalMarketing, "is AEO actually worth it on top of SEO"
69 upvotes, 131 comments, 2026-03-09. A buyer asks their own agency's upsell pitch back to a wider audience and gets a mixed, skeptical thread rather than a confirmation.
r/seogrowth, "is it worth focusing on GEO right now"
53 upvotes, 56 comments, 2026-03-04. Practitioners discuss whether generative-engine optimization spend is defensible before a client has asked for it.
LinkedIn, "spent months hunting for one AI visibility tool that does everything"
Greg Bennett, 2026-08-03. A buyer's own account of evaluating the monitoring-tool category and never finding one that answered the causation question, only the counting one.
LinkedIn, "we spent 6 months testing every AI visibility tool on the market"
Sebasttian Pinzon, 2025-11-20. Names four monitoring vendors tested end to end, evidence that the category's own buyers already do comparison homework the ranked listicles claim to replace.
Gary Vaynerchuk, on agencies building GEO/AEO content as a standing line item
193 favorites, 21 retweets, 75,125 views, 2026-07-27. Category-momentum evidence that AEO/GEO is now a budget line agencies build around, independent of whether any single agency can prove its own effect.
Our own step-zero measurement run
12 money queries times 5 identical repeats, one model, one day, 60 of 60 calls succeeded, 7 of 12 flipped outcome between identical repeats. Published with its numbers on this site.
Our own permutation test
20,000 random splits of the same query set with no intervention applied. Null difference centred on zero, mean +/-0.0016, standard deviation 0.215.
SparkToro and Gumshoe, repeated-prompt consistency
Under a 1 in 100 chance that two identical prompts return the same brand list. Independent evidence of the same instability our own repeat run measured, on a different corpus.
Atil et al., nondeterminism at temperature zero
Up to 15 percent accuracy variation across ten runs at temperature zero, across five models and eight tasks. The instability a single-reading score hides is an infrastructure property, not a setting anyone forgot to turn off.
Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments
The canonical text on why a held-out control is the correct instrument for a noisy metric, and on sizing the minimum detectable effect before trusting a result.
Backlinko, category framing of the monitoring-tool market
One of ten ranking guides for "ai visibility tool," read for how the category itself talks about what these tools do and do not claim to measure.

Questions this answers

What does an AI visibility tool actually measure?
It counts how often a brand name is returned when a set of queries is asked of one or more AI systems, then reports that count as a score. It is a measurement of appearance frequency, not a measurement of what caused that frequency to change between two readings.
Can a rising AI visibility score prove an AEO agency's work is effective?
Not on its own. A rising score is compatible with real work causing it, with the underlying model changing its own answers, or with both happening at once. Separating the three requires a comparison point the score alone does not provide.
What is a held-out control in AEO measurement?
A set of queries deliberately left unworked for the same period as the queries being actively optimized. Comparing the two sets at the end of the window lets whatever the system did on its own subtract out of the reported difference, isolating the attributable part.
Why don't AI visibility dashboards run a held-out control by default?
A control arm costs a real, deliberately unworked slice of query coverage for the whole measurement period, and it delays any reportable number until the window closes. A single score ships immediately and looks more decisive on a slide, even though it answers a narrower question.
How many repeated queries does it take to trust a visibility reading?
It depends on the effect size you expect. A design testing twelve queries five times each has a minimum detectable difference above fifty percentage points, far too coarse for most real engagements. A larger query set with more repeats brings that floor down to a level closer to what an actual AEO engagement moves.
What should I ask an AEO vendor before trusting their reported number?
Ask what the number would have read with nobody working the account. Ask how many times each query was repeated. Ask whether the reported figure is a level or a difference against a control. A vendor with a real answer to all three has done the work this post describes.

Keep reading