Measuring lift, answer engine work
A chart going up is not evidence that the work you paid for did it.
Four kinds of evidence get offered for answer engine work, and three of them cannot separate what you paid for from what your category did anyway. Here is what each one settles, what it leaves open, and how to run the fourth on your own queries. Everything on this page is runnable without us.
Three of the four fall short of the same line, and they fall short of it for the same reason rather than for three different ones.
9.9pp
smallest lift the instrument can detect
At 40 queries by 40 samples. Anything smaller is unproven.
45.5%
of citations change between consecutive observations
Ahrefs, 43,000 keywords. AI Overviews persist 2.15 days.
7 of 12
queries flipped outcome across identical repeats
One model, five repeats each. A single read is a coin flip.
The short version
One definition, in a paragraph a model can quote.
Definition first, under sixty words, no pronoun pointing at anything outside itself. It is the same shape we build for clients, on our own page, which is the only honest way to sell it.
Short answer
How do you measure whether answer engine optimization worked?
Causal-lift measurement is the practice of reading answer engine results against an arm deliberately left alone, rather than against last quarter. A treated set of queries and a matched held-out set are written down before any work starts, both are sampled at day zero and day ninety, and the difference between their two changes is the number.
Causal-lift measurement is the practice of reading answer engine results against an arm deliberately left alone, rather than against last quarter. A treated set of queries and a matched held-out set are written down before any work starts, both are sampled at day zero and day ninety, and the difference between their two changes is the number.
Causal-lift measurement is the practice of reading answer engine results against an arm deliberately left alone, rather than against last quarter.1
Same words, both times. A paragraph that has to be summarised before it can be used is a paragraph an assistant paraphrases and does not cite.
The four rungs
Four rungs, and one of them answers the question you are asking.
Sorted by what each one can settle, not by how convincing it looks in a deck. The first three are cheap, fast and sincere, and none of them can tell your movement apart from your category's.
One column is full and the other has a single mark in it. That ratio is the argument, and it does not change with how much anybody spends.
The screenshot
Leaves it open
- one answer, one moment
- pasted into a channel
- prompt not recorded
- What it decides
- That the sentence existed once, on one engine, for whoever ran it. The prompt, the account and the day are all uncontrolled, so nothing about it predicts what an assistant tells the next buyer who asks.
- Who holds it today
- Everybody. It is the cheapest artifact in this category and the most circulated, and it is what a vendor reaches for in the first week because it is the only thing available that fast.
The visibility score
Leaves it open
- one sweep, one number
- one run per query
- no repeat, no interval
- What it decides
- How often you were named in one sweep. Not whether the same sweep would return the same number an hour later: we ran twelve money queries five times each on one model on the same day, changed nothing between runs, and seven of the twelve flipped outcome.
- Who holds it today
- Every monitoring product in the category, which is why this part is handed over free. A number that cannot survive its own repeat is not a thing to charge for.
The before and after
Leaves it open
- same queries, two dates
- day 0 against day 90
- nothing held back
- What it decides
- That the level moved. Not who moved it. Citation sets churn without anybody's help: Ahrefs measured 45.5 percent of citations changing between consecutive observations across 43,000 keywords, so a ninety-day difference contains your work and the category's drift added together and inseparable.
- Who holds it today
- Most agencies in this category, and it is a genuine step up from a screenshot. It is also the last rung where the vendor grading the work is the only party who can check it.
The held-out control arm
Settles it
- treated arm and control arm
- split written down at day 0
- difference in differences
- What it decides
- How much of the movement the work caused. Unbiased noise cancels between two arms: across 20,000 random splits of our own data with no work applied to either side, the null difference centred on zero, which is the finding the whole design rests on.
- Who holds it today
- Nobody we have found. Across the 64 companies in our own audit, not one published a control group or a causal-lift number, and that empty square is the reason this page exists.
Three of the four are honest and none of them can answer the question you are paying to have answered. The fourth can, and the price of it is ninety days and a set of queries you agree not to touch.
One case, all the way through
The verdict falls out of the trace, not out of an opinion.
One claim, taken apart. This is the shape of the number most vendors in this category report at day ninety, and what is left of it once you ask what it was measured against.
citations up 41 percent in ninety days
- The claimCitation share on the tracked query set read 41 percent higher at day ninety than at day zero.
- The sampleOne run per query per date. Seven of our own first twelve queries changed outcome across five identical repeats on a single day, so one run per date is a coin flip at both ends.
- The controlNothing held back. Every query in the set was worked, so there is no arm the category's own drift can be read off and subtracted.
- The thresholdRecorded after the reading. A target chosen once the number is known cannot be missed, which makes the result a description rather than a test.
- What survivesThat the level is higher than it was. Ahrefs measured 45.5 percent of citations changing between consecutive observations, so a level that moved is the expected state rather than the finding.
Not distinguishable from a category that moved on its own
The verdict is read off the column of marks. Nothing in the trace needs an opinion about the vendor to reach it.
- The claim
- A level, measured twice, with nothing held back
- The sample
- One run per query, on a set that flips on repeat
- The threshold
- Chosen after the number was already known
- What it settles
- That the level changed. Not who changed it
Run it yourself
Five minutes settles whether any of this is worth paying for.
Four questions, asked of any vendor in this category and of us. They take five minutes on a call and they settle whether anybody in the room can tell a result from a coincidence.
- 01
What is held out
Ask which of your queries are being deliberately left alone for the whole engagement. If the answer is none, every number you will be shown is a level, and a level moves on its own.
- 02
When the threshold was written
Ask for the date the target was recorded, not the target. A threshold written after the reading is a description of what happened, and it cannot be missed.
- 03
How many repeats per query
Ask how many times each query is asked before it counts as measured. We run five, because seven of our first twelve queries changed outcome across five identical repeats on one day.
- 04
What the untouched arm did
Ask what the held-out queries did over the same ninety days. Without that number nobody in the room can separate your movement from the category's.
All four answered
Somebody is already running a control arm on your account. There is nothing here for us to sell you, and the right next move is to ask everyone else pitching you the same four questions.
Some answered
You are measuring, and the measurement cannot settle the argument in either direction yet. That is the position this work is built for, because the gap is real and it is closable.
None answered
Whatever you are being shown is a level reading with a chart around it. The first job is a design rather than more work, because more work measured this way produces the same unfalsifiable number faster.
A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.
Four questions, asked in one sitting, answerable by anybody being paid to do this work. An answer that needs a follow-up meeting is an answer.
A perfectly on-vertical category definer is a bad engagement for both sides. A slightly off-vertical challenger is a good one.
The method, given away
How to run it without hiring anyone.
The design is not proprietary and running it badly is worse than not running it at all. This is the same sequence we run, written out so you can run it on your own queries.
- 01Freeze the query set
- 02Split it at random
- 03Write the threshold downFixed first
- 04Sample both arms five times
- 05Take the difference of the differences
Four of the five can be reordered without much cost. The one marked cannot be moved at all once anything else has started.
- 01
Freeze the query set
Take twenty to thirty questions with a purchase behind them and stop editing the list. A set that grows during the window cannot be compared with itself, and every addition is a query chosen because of how it looked that week.
- 02
Split it at random
Divide the set in two by coin flip. One half gets worked, the other half is deliberately left alone for ninety days, and neither half is picked for how it reads today.
- 03
Write the threshold down
Record what result would count as a win and the date you recorded it, before any work starts. This is the step that gets skipped, and it is the one that makes every number after it mean something.
- 04
Sample both arms five times
Ask every query in both arms five times on the same engine at day zero, and again at day ninety. One read decides nothing, which is measured rather than asserted: seven of our first twelve flipped.
- 05
Take the difference of the differences
Subtract the control arm's change from the treated arm's change. What is left is the part the work is responsible for, and it is the part nobody can hand back to the category.
What counts
Most of what gets offered here does not survive the first question.
Most of what is presented as evidence in this category is a level reading with a chart drawn around it. The cuts are what leave you holding something that can settle an argument.
everything anybody offers you
The counts are the two columns below this, not an estimate of what a typical intake survives. Nothing here claims a rate we have not measured.
Keep it if
- A number read off an arm that was deliberately left alone for the whole window.
- A threshold recorded before the work started, with the date it was recorded on it.
- A query set frozen at day zero and sampled the same way at both ends.
- A stated repeat count, with the flip rate that made the repeats necessary.
Cut it if
- A screenshot of one answer. It settles that the sentence existed once, for one person.
- A visibility score from a single sweep. Twelve of ours were run five times each and seven changed outcome.
- A before and after with nothing held back. The category moves on its own and you cannot subtract it.
- A target chosen once the number was known, which is a description wearing the clothes of a test.
What it is measured against
The number is a difference between two differences.
A design is not a report. Once the split is written down, a list of queries becomes the thing every later claim is measured against, including ours, and that is the whole product.
- Validated
- 5 identical repeats before one query counts
- Split
- A treated arm and a held-out control arm
- Registered
- Both written down before any work starts
- Depth
- 40 queries by 40 samples, 9.9pp floor
- Readout
- Difference-in-differences at day 90
A single rising line cannot draw this, because the quantity it is missing is the one the untouched arm was holding.
Ninety days is the shortest window in which this number can exist, which is why every tier carries a 90-day minimum. Below it the treated and held-out arms have not had time to diverge, so there is nothing to read whatever anybody is paying. Diagnosis lands inside the first three weeks; the causal number lands at day ninety.
Start here, free
We run it on your queries: 12 queries, 5 repeats each, 3 engines.
Delivered in five working days, and walked through live rather than emailed as a PDF. You get your questions built and disambiguated, your failure mode named, the hosts cited on those questions ranked by citation depth, and a pre-registered baseline you can hold anyone to afterwards, including us.
The rest of the cluster
One method, two surfaces, one number at the end of both.
One instrument, pointed at two surfaces. The guide you are reading is the free half of whichever one your category actually turns on.
Questions
The five we get asked about this.
Not covered here? The Gap Report costs nothing and answers most of the rest with your own data.
Why not just compare our citations before and after?
A before and after tells you the level moved. It cannot tell you who moved it, because the level moves on its own: Ahrefs measured 45.5 percent of citations changing between consecutive observations across 43,000 keywords, while the answer stayed 95 percent semantically identical. A held-out arm is what separates your movement from the category's, and nothing else does.
How small a lift can this design actually detect?
At 40 queries by 40 samples the smallest difference it can tell apart from zero is 9.9 percentage points. Thinner designs see less, which is why the entry tier reports the causal number without guaranteeing it: at 20 to 30 queries the minimum detectable effect is 11.4 to 14 points, larger than a plausible real win, so a genuine result would come back as not distinguishable from zero.
Does holding queries back mean we get less work?
The control arm is a fraction of the set and it is chosen at random rather than by value, so nothing is being withheld on the basis of what it is worth. What the arm buys is the ability to say how much of the movement was yours. Across the 64 companies in our audit, not one published that number.
How long before there is a number?
Ninety days, which is why every tier carries a 90-day minimum. Diagnosis lands inside the first three weeks and the causal-lift report lands at day ninety, because that is the shortest window in which a treated arm and a held-out arm have had time to diverge enough to read.
What is in the free Gap Report?
Twelve money queries, five repeats each, across three engines, delivered in five working days. It names which of the nine failure modes you are in, ranks the hosts cited on your queries by citation depth, and sets a pre-registered baseline. It is walked through on a call rather than emailed as a PDF.
Free gap report
When the model answers,
be the one it names
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 3 engines · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site