New: the nine ways developer tools go invisible in AI answers.
citon

Answers

How do you prove AEO work caused the AI citation lift, not just coincidence?

Short answer

How do you prove AEO work caused the AI citation lift, not just coincidence?

With a pre-registered, held-out control arm and enough queries to clear a measured noise floor, not with a single before-and-after reading. Assistant answers are unstable between identical asks, our own step-zero run found 7 of 12 queries changed outcome across five identical repeats, so a raw before-and-after comparison cannot separate real movement from the model changing its own mind. The fix is a difference-in-differences design: a treated group that gets the work and a control group that does not, both measured at the same two timepoints. Our own permutation test over 20,000 random splits with zero intervention applied found the null difference centers on a mean of roughly plus or minus 0.0016 with a standard deviation of 0.215, which means unworked noise cancels across a large enough sample and a real lift stands out from it.

The short version

Before-and-after readingNot proof
Held-out control armIsolates the effect
40 x 40 designDetects 9.9pp

Null-case check across 20,000 splits: unworked noise centers on roughly zero.

01

Why a before-and-after reading is not proof

The most common way AEO results get reported is a visibility score before an engagement and the same score after, presented as the lift. That comparison has no way to separate work-caused movement from the model itself changing its answers between the two reads, because nothing in a single before-and-after design accounts for the underlying instability. Our own instrument check found 58.3 percent of tracked queries, 7 of 12, returned a different outcome across five identical repeats with nothing else changed.

A number that moves for reasons unrelated to the work looks identical to a number that moves because of it. Without a comparison point that isolates the work, a rising score and a falling one are both compatible with having done nothing at all.

20,000 splits, zero intervention applied

Null mean difference+/- 0.0016
Standard deviation0.215
ReadingUnbiased, cancels at scale

Confirming the untreated case centers on zero is what makes a difference-in-differences result trustworthy.

02

The design that actually isolates the effect

A difference-in-differences design fixes this by adding a second, deliberately untouched arm: the same class of queries, measured at the same two timepoints, with no placement work applied. Whatever the untouched arm moves by is the model's own churn for that window, and subtracting it from the treated arm's movement leaves the part attributable to the work.

We validated the null case directly before trusting this design on real work: running 20,000 random splits with no intervention applied at all, the resulting difference centered on a mean of roughly plus or minus 0.0016 with a standard deviation of 0.215. That is the honest finding, the unworked noise is unbiased and it cancels at scale, which is exactly the property a difference-in-differences design needs to be trustworthy.

Difference-in-differences, two timepoints

Treated arm, timepoint 1 -> 2Work applied
Control arm, timepoint 1 -> 2Left untouched

The control arm's movement is subtracted from the treated arm's. What is left is attributable to the work.

03

What sample size the proof actually needs

Isolating a real effect from that noise floor costs samples, and the cost is steep at small scale. Computed from the variance in our own run, a 12-query by 5-repeat design can only detect a change of 51.1 percentage points, which is unusably coarse. A 20 by 20 design reaches 21.4 points, and a 40 by 40 design, roughly 3,200 calls and about six hours per timepoint, reaches 9.9 points.

That is the honest price of a provable answer rather than an asserted one. A vendor quoting a lift number from a small query set is either reporting something too coarse to mean anything or has not run the arithmetic that would tell them so.

Minimum detectable lift, our own design

12 x 5 queries x repeats51.1pp
20 x 20 queries x repeats21.4pp
40 x 40 queries x repeats9.9pp

A vendor quoting a lift number from a small query set is reporting something too coarse to mean anything.

Asked next

The questions that follow this one

Is a difference-in-differences design overkill for a small business?

It scales down, not away. The design principle, treated versus untouched, measured at two timepoints, holds at any size; what shrinks is the minimum detectable lift, which is why our entry tier reports the causal-lift number without attaching a guarantee to it.

Why does the noise cancel instead of biasing the result?

Because across 20,000 random splits with no work applied, the null difference centered on roughly zero rather than drifting positive or negative. An unbiased noise source averages out as the sample grows, which is the statistical property the whole design depends on.

Can I run this measurement myself before hiring anyone?

Yes, the design is not proprietary. Pick a query set, split it into a treated and control half, apply work to only one half, and measure both at two timepoints. The hard part in practice is discipline: not touching the control arm, and running enough queries to clear the noise floor.

Why is there no guarantee below the 10,000 dollar tier?

Instrument power, not confidence. At 20 to 30 queries the minimum detectable effect is 11.4 to 14 percentage points against a plausible real lift of 5 to 15, so a genuine win could read as statistically indistinguishable from zero. The causal-lift number is still reported at every tier.

Building your own design

1Pick a money-query set and split it into a treated and control half.
2Apply work only to the treated half for the full test window.
3Measure both halves at two timepoints, not one.
4Size the query count to the lift you actually expect to see.

Free gap report

When the model answers,
be the one it names

Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.

12 queries · 5 repeats each · 5 working days

Free · Walked through live · Five working days

Gap ReportSample
12 queries · 5 repeats each

Queries we run

best rate limiting api
your-api alternativesYou, 1 of 12
cheapest webhook api
api gateway for startups

Cited instead of you

reddit.com38%
g2.com26%
news.ycombinator.com21%

91.5% of citations point off-site

Your failure mode, named
Hosts ranked by citation depth
The pre-registered baseline