Answers
How do you prove AEO work caused the AI citation lift, not just coincidence?
Short answer
How do you prove AEO work caused the AI citation lift, not just coincidence?
With a pre-registered, held-out control arm and enough queries to clear a measured noise floor, not with a single before-and-after reading. Assistant answers are unstable between identical asks, our own step-zero run found 7 of 12 queries changed outcome across five identical repeats, so a raw before-and-after comparison cannot separate real movement from the model changing its own mind. The fix is a difference-in-differences design: a treated group that gets the work and a control group that does not, both measured at the same two timepoints. Our own permutation test over 20,000 random splits with zero intervention applied found the null difference centers on a mean of roughly plus or minus 0.0016 with a standard deviation of 0.215, which means unworked noise cancels across a large enough sample and a real lift stands out from it.
The short version
Null-case check across 20,000 splits: unworked noise centers on roughly zero.
01
Why a before-and-after reading is not proof
The most common way AEO results get reported is a visibility score before an engagement and the same score after, presented as the lift. That comparison has no way to separate work-caused movement from the model itself changing its answers between the two reads, because nothing in a single before-and-after design accounts for the underlying instability. Our own instrument check found 58.3 percent of tracked queries, 7 of 12, returned a different outcome across five identical repeats with nothing else changed.
A number that moves for reasons unrelated to the work looks identical to a number that moves because of it. Without a comparison point that isolates the work, a rising score and a falling one are both compatible with having done nothing at all.
20,000 splits, zero intervention applied
Confirming the untreated case centers on zero is what makes a difference-in-differences result trustworthy.
02
The design that actually isolates the effect
A difference-in-differences design fixes this by adding a second, deliberately untouched arm: the same class of queries, measured at the same two timepoints, with no placement work applied. Whatever the untouched arm moves by is the model's own churn for that window, and subtracting it from the treated arm's movement leaves the part attributable to the work.
We validated the null case directly before trusting this design on real work: running 20,000 random splits with no intervention applied at all, the resulting difference centered on a mean of roughly plus or minus 0.0016 with a standard deviation of 0.215. That is the honest finding, the unworked noise is unbiased and it cancels at scale, which is exactly the property a difference-in-differences design needs to be trustworthy.
Difference-in-differences, two timepoints
The control arm's movement is subtracted from the treated arm's. What is left is attributable to the work.
03
What sample size the proof actually needs
Isolating a real effect from that noise floor costs samples, and the cost is steep at small scale. Computed from the variance in our own run, a 12-query by 5-repeat design can only detect a change of 51.1 percentage points, which is unusably coarse. A 20 by 20 design reaches 21.4 points, and a 40 by 40 design, roughly 3,200 calls and about six hours per timepoint, reaches 9.9 points.
That is the honest price of a provable answer rather than an asserted one. A vendor quoting a lift number from a small query set is either reporting something too coarse to mean anything or has not run the arithmetic that would tell them so.
Minimum detectable lift, our own design
A vendor quoting a lift number from a small query set is reporting something too coarse to mean anything.
Asked next
The questions that follow this one
Is a difference-in-differences design overkill for a small business?
It scales down, not away. The design principle, treated versus untouched, measured at two timepoints, holds at any size; what shrinks is the minimum detectable lift, which is why our entry tier reports the causal-lift number without attaching a guarantee to it.
Why does the noise cancel instead of biasing the result?
Because across 20,000 random splits with no work applied, the null difference centered on roughly zero rather than drifting positive or negative. An unbiased noise source averages out as the sample grows, which is the statistical property the whole design depends on.
Can I run this measurement myself before hiring anyone?
Yes, the design is not proprietary. Pick a query set, split it into a treated and control half, apply work to only one half, and measure both at two timepoints. The hard part in practice is discipline: not touching the control arm, and running enough queries to clear the noise floor.
Why is there no guarantee below the 10,000 dollar tier?
Instrument power, not confidence. At 20 to 30 queries the minimum detectable effect is 11.4 to 14 percentage points against a plausible real lift of 5 to 15, so a genuine win could read as statistically indistinguishable from zero. The causal-lift number is still reported at every tier.
Building your own design
Free gap report
When the model answers,
be the one it names
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site