Answers
How long does answer engine optimization take to work?
Short answer
How long does answer engine optimization take to work?
Two different clocks, and conflating them is how this question gets answered badly. Citations themselves turn over fast: Ahrefs, measuring across 43,000 keywords, found AI Overviews persist about 2.15 days and 45.5 percent of citations change between consecutive observations, so the sources behind an answer can move within a week of work landing. Proving that the movement was caused by the work is the slow clock, and it needs a held-out control group plus enough sampling to clear the noise, which is why every engagement here runs a minimum of 90 days. Anyone quoting you a faster timeline to a proven result is quoting the fast clock and calling it the slow one.
Two clocks, not one
Rotation measured by Ahrefs across 43,000 keywords. The 90 days is our own minimum term, and it is a statistics bound rather than an effort one.
01
The fast clock: citations rotate in days
The sources behind an AI answer are far less stable than most people assume. Ahrefs measured this across 43,000 keywords and found AI Overviews persisting roughly 2.15 days, with 45.5 percent of citations changing between consecutive observations, while the answer itself stayed about 95 percent semantically identical. The model keeps saying the same thing and keeps changing who it credits for saying it.
That is genuinely good news for a challenger. A citation set that rotates every few days is one you can enter without waiting for an incumbent to decline. It is also why early movement is a weak signal, because a set that volatile moves on its own.
Ahrefs, 43,000 keywords
2.15
days an overview persists
45.5%
of citations change
Measured by Ahrefs at source, between consecutive observations, while the answer itself stayed about 95 percent semantically identical.
02
The slow clock: proving it was you
Because the citation set moves by itself, seeing your name appear after doing some work tells you almost nothing about whether the work caused it. Separating the two needs a comparison against queries you deliberately did not touch, measured on the same schedule, so that whatever moved everything is visible in both and whatever moved only the treated set is attributable.
This is the part the category tends not to publish, and it is the reason for a 90-day minimum on every engagement here rather than a shorter pilot. A control group is not a reporting nicety, it is the only thing standing between a real result and a coincidence you have paid for.
Why one arm proves nothing
If both arms move, the churn moved them. Only the difference between them is attributable to the work, which is why the second arm is not optional.
03
What the sampling actually costs
The timeline is also bounded by how much measurement is needed for a result to be distinguishable from zero. On our own numbers a 12 query by 5 repeat design has a minimum detectable lift of 51.1 percentage points, which will never show a realistic change. Getting that to 9.9 points takes roughly 40 queries by 40 samples, about 3,200 calls and around six hours of measurement per timepoint.
Two timepoints and a control arm is the smallest honest design, and it is why the answer to how long this takes is a function of statistics rather than of effort. You can work faster. You cannot detect faster.
Detection floor against a plausible result
A real change smaller than the floor comes back as no change at all. The first design cannot see the third row.
Asked next
The questions that follow this one
Can I see movement in the first month?
Often, because citation sets rotate on the order of days. What you cannot do in the first month is tell whether the movement was yours, which is a different claim and needs a control arm to support.
Why is the minimum term 90 days?
Because a causal read needs a held-out control measured across enough timepoints to separate a real effect from the churn. Shorter engagements can do work and cannot prove it, and reporting a number we cannot stand behind is worse than reporting none.
Why not just check weekly and watch the trend?
Single readings flip. In our step-zero run 7 of 12 queries returned different outcomes across five identical repeats, so a weekly line chart of a small sample is mostly a picture of noise with a trend the eye supplies.
Do you guarantee a result?
The causal lift number is reported at every tier. A money-back guarantee attaches only from the $10,000 tier upward, because below that the query volume makes the instrument too underpowered to hold us to a threshold fairly.
What a real 90 days looks like
Free gap report
When the model answers,
be the one it names
Tell us your category and the questions your buyers ask. We run them against live AI answers and walk you through what came back. If you are already winning, we will tell you that too.
12 queries · 5 repeats each · 5 working days
Free · Walked through live · Five working days
Queries we run
Cited instead of you
91.5% of citations point off-site