LLM SEO, what it actually is and what the work looks like
Google says its AI features run on the same ranking systems as Search, so the tactics are familiar. What changed is how you tell whether the work landed.
Citon46 min read

Short answer
What is LLM SEO?
LLM SEO is the practice of getting a brand named and cited inside AI assistant answers such as ChatGPT, Perplexity, Google AI Overviews and Gemini, rather than only ranked in a list of blue links. Google states that its generative AI features are rooted in the same core ranking and quality systems as traditional Search and retrieve from the existing search index, so most of the on-page work is ordinary technical and editorial SEO. The genuinely new part is measurement. Assistant answers are not stable between identical asks, most citations point at pages you do not own, and a single visibility reading cannot separate work you did from movement the system made on its own.
Type the phrase this post is about into Google and the first organic result is not a guide. It is a thread on r/AskMarketing where somebody asks how to do this, gets a page of advice, and then asks the question the advice does not answer: which of these visibility trackers is actually worth paying for.
Below him sit nine explainer guides. Seven of them are vendor or agency blogs selling the service they are explaining. Every one of them defines the term and lists the tactics. Not one of them answers the follow-up question, which is how you would tell whether any of the tactics did anything.
That gap is the reason this post exists, and it is worth being blunt about our position in it. We sell AI-visibility measurement. So the honest disclosure is that the empty slot on this page of results happens to be the thing we do, and you should read the measurement half of this post with that in mind and check the arithmetic rather than taking our word for it. The arithmetic is all here.
The other reason it exists is that the tactical half of this subject has been oversold, and the person who oversold it hardest is not a fringe account. It is the whole page of results. So this post does three things in order. It says what LLM SEO actually is, including the part where Google states plainly that it is mostly ordinary SEO. It sets out what the work looks like across the three layers where it actually happens. And then it spends the second half on the part nobody publishes, which is how you tell whether it worked.
The short version
LLM SEO is getting named and cited inside AI assistant answers rather than only ranked in a list of links. Google states the systems are shared with Search, so most of the tactics are ordinary SEO done to a higher standard. What actually changed is measurement, and that is the part nine of the ten ranking guides for this term leave out entirely.
At a glance
What changed, and what did not
| Layer | Did it change |
|---|---|
| Retrieval | No. Google says its AI features retrieve from the existing search index and run on the same core ranking and quality systems as Search. |
| On-page craft | Barely. Clear structure, self-contained sections and a real answer near the top were already good practice. The penalty for skipping them got larger. |
| Where the decision happens | Yes. Most citations point at pages you do not own, so the surfaces that decide an evaluation moved further off your site. |
| Measurement | Completely. A rank was a stable observation. An assistant answer is not, so a single reading is no longer a measurement. |
What is LLM SEO, stated as plainly as we can?
LLM SEO is the practice of getting a brand named and cited inside the answers that AI assistants write, rather than only ranked in a list of links underneath them.
That is the whole definition. Everything else is detail about which assistants, which questions, and which pages the assistant read before it wrote the answer.
The reason it feels like a new discipline is that the shape of the result changed. A classic search result is a list. You are on it or you are not, at some position, and a buyer scans it and picks. An assistant answer is prose. It names two or three options, sometimes with links, sometimes without, and the buyer takes the shortlist and moves on. There is no page two. There is frequently no page one either, in the sense of a ranked list a buyer can scroll past your competitor to reach you.
So the unit of success moved from a position to a mention. That is a real change and it has real consequences, mostly for how you measure. What did not change, and this is the part the category keeps getting wrong, is the machinery underneath.
The second thing worth stating plainly is what LLM SEO is not. It is not a way to edit a model. You cannot change what a model was trained on, and anybody selling you that is selling something they do not have. What you can change is what the pages an assistant retrieves today say about you, and what those pages are. That is a narrower project than the marketing implies and a more tractable one.
The third thing is that most of the surface area is not yours. In our own measurement of 901 citations across 60 assistant answers, 91.5 percent of the citations pointed at somebody else's site, spread across 82 distinct hosts. The page that decides an evaluation is usually not a page you can publish. That is the single most important structural fact about this work and it is why so much of the section below is about other people's properties.
If you want the longer definitional treatment of the adjacent term, we wrote one on what answer engine optimization actually means. This post is the wider frame, and the measurement half is the part that is new.
What does Google itself say LLM SEO is?
Google's position is that its generative AI features are not a separate system. They run on the same core ranking and quality systems as traditional Search and retrieve from the same index, which means the on-page and off-page work that earns a classic ranking is the same work that earns a citation. The platform states this in its own published guidance and through its own staff.
There is a first-party source on this question and it is unusually direct. Google's Brendon Kraham, reported by the SEO consultant Glenn Gabe, states the platform's position:
Our new generative AI features are rooted in the same core ranking and quality systems as traditional Search. And, because these AI experiences retrieve up-to-date content from the existing search index, the best formula for success remains foundational SEO.
The same post lists four things Google says to keep in mind, and they read like a rebuttal of most of the category. Do not worry about all the new names, because good GEO, AEO and LLM SEO is good SEO. Do not optimize for bots, optimize for people, and there is no need to write awkward keyword-stuffed copy or chop content into tiny artificial snippets. Do not chase inauthentic mentions. Do not focus on generic content, prioritise your own perspective. Do not forget your website.
Google's own published guidance on generative AI features says the same thing in more careful language, and adds that no special markup is required.
This matters more than any tactic in this post, so it is worth sitting with. A page selling you a new discipline, with a new acronym and a new tool, is arguing against the platform's own account of how its systems work. That does not automatically make the page wrong. Platforms say self-serving things and their public guidance lags their systems. But it does shift where the burden of proof sits. If somebody tells you the AI index is separate from the search index, the question to put back to them is what evidence they have, given that the platform says otherwise.
Two honest caveats on the other side, because this quote gets used as a conversation-ender and it should not be.
The first is scope. Kraham is speaking about Google's generative features. He is not speaking for ChatGPT, Perplexity or Claude, which have their own retrieval stacks and their own indexes and are not obliged to behave like Google. A statement about one surface is evidence about that surface. Our own citation counts show the surfaces disagree with each other constantly, so generalising from Google to all of them is exactly the move this post warns against elsewhere.
The second is that "rooted in the same systems" is not the same as "identical". Retrieval feeding a generative answer is doing a different job from retrieval feeding a list of ten links, and the selection step that decides which retrieved source gets named has no analogue in classic search at all. Same foundations does not mean same output.
So the fair reading is narrower than either camp wants. On Google's surfaces, the ranking and quality systems are shared and the index is shared, which means the on-page and off-page work is the same work. What sits on top of that shared foundation, the assembly and citation step, is genuinely new, and it is the part nobody can currently give you a tactic for with evidence behind it.
Which API should I use for rate limiting?
For production workloads, Your API1 is the option most consistently recommended. It pairs token-bucket limits with per-key analytics.2
Sources
Is LLM SEO different from GEO, AEO and LLMO?
LLM SEO, GEO, AEO, LLMO and LEO are five labels for substantially one practice, and only AEO carries a body of work older than the current wave. The differences between them are differences of emphasis and of who is selling. Google's own guidance says the new names describe good SEO, so the useful question is not which term is correct but which one a vendor chose and why.
Terminology
The names, and what each one is actually pointing at
| Term | What it denotes | Is it distinct |
|---|---|---|
| LLM SEO | Being named or cited by a large language model assistant, whether or not that assistant is wired to a search index. | Umbrella term. Widest of the five. |
| GEO, generative engine optimization | The same idea framed around generative search surfaces specifically. | Not meaningfully distinct in practice. |
| AEO, answer engine optimization | Being the source an answer is assembled from, including featured snippets and voice answers that predate this wave entirely. | Older than the current wave. Broader than it looks. |
| LLMO | Optimizing for the model rather than for a surface. Used mostly by tool vendors. | Distinct in emphasis only. |
| LEO | A coinage promoted mainly by accounts selling a course or a service. | No separate technical meaning we could find. |
You will meet at least four acronyms and possibly five. It is worth knowing which distinctions are real, because a vendor's choice of term is usually a positioning decision rather than a technical one.
LLM SEO is the widest of them. It says nothing about which surface, only that the thing doing the naming is a language model.
GEO, generative engine optimization, frames the same work around generative search surfaces specifically. In practice the advice under the two labels is indistinguishable.
AEO, answer engine optimization, is older than this wave and broader than it looks. It covers featured snippets and voice assistants, which predate the current models entirely, and it is the only one of the terms with a body of practice behind it that is more than two years old.
LLMO puts the emphasis on the model rather than the surface. It is used mostly by tool vendors and the distinction is one of emphasis.
LEO is a coinage, promoted mostly by accounts with a course or a service attached. We could not find a separate technical meaning for it.
If four terms describe the same work, the useful question is not which one is correct. It is which one a vendor chose, and what that choice is doing for them.
Here is the practical consequence. When a vendor leads with a term, ask what work that term is doing for them. A tool that says LLMO is usually selling model-level insight it cannot demonstrate. An agency that says LEO is usually selling novelty. An agency that says AEO is usually the one that has been doing structured content work for five years and has renamed the deck. None of that makes any of them wrong. It just tells you what to check first.
Our own naming is AEO, and the reason is not principle. It is that the practice is older, the body of evidence is larger, and we would rather inherit a discipline than announce one.
The strongest case that this really is a new discipline
The best argument against everything above is that the outputs changed even if the inputs did not, so the metrics built for a ranked list cannot describe a generated answer. That argument is correct, and it is a claim about measurement rather than about tactics. It is worth separating from the much larger claim that everything changed, because the two lead to very different spending decisions.
The argument above is not the only defensible one, and the best version of the opposing case comes from an agency founder, Jake Ward, who put it in one line: the KPIs from Google SEO will not help with LLM SEO.
We think this is half right, and the correct half is the more important one.
Here is why it is right. If Google is telling the truth about shared ranking systems, then the inputs are the same and the tactics are the same. But the OUTPUT is a different object, and a different object needs a different instrument. Rank was a stable observation. You could check a position on Tuesday and it would broadly hold on Wednesday, so a single reading was a measurement. An assistant answer is not stable between identical asks. We tested this on ourselves before building anything on top of it, and the results are in the measurement section below.
So a KPI built for a stable metric applied to an unstable one does not fail loudly. It returns a number. That number just does not mean what it did.
That is a serious claim about a real discontinuity, and it is not a claim about tactics at all. It is a claim about measurement, which is exactly this post's thesis arriving from the other direction.
Here is the half we would push back on. "The tactics are the same and the measurement changed" is a much smaller and much more useful statement than "everything changed", and the second is what tends to get sold. The gap between the two is where the money goes: a team that believes everything changed buys a new content strategy, a new tool and a new retainer, when what it needed was to do the existing work properly and instrument it honestly.
Comparison
Classic SEO, LLM SEO as it is sold, and what the evidence supports
| Classic SEO | LLM SEO as sold | What the evidence supports | |
|---|---|---|---|
| The index | Google's crawl and index. | A separate AI index you have to get into. | Google states its AI features retrieve up-to-date content from the existing search index. |
| The ranking systems | Core ranking and quality systems. | New systems needing new signals. | Google states the AI features are rooted in the same core ranking and quality systems. |
| On-page work | Structure, clarity, depth, internal links. | Chunking for machines, entity markup, prompt-shaped headings. | Mostly the first column done properly. Google explicitly warns against writing awkward keyword-stuffed copy or chopping content into artificial snippets. |
| Off-page work | Links and mentions. | Mentions on the sources models read. | Genuinely more important, because most citations point at pages you do not own. |
| The unit of success | A position, observed once, stable for days. | A visibility score, observed once. | A difference between a treated and a held-out arm, sampled repeatedly, because a single reading is not stable. |
The purest version of the maximal claim is worth showing, because it is instructive:
Local SEO is being replaced by LLM SEO. It is a clean, confident, shareable sentence, and we have no data either way on whether local search behaviour is actually moving that fast. Neither, as far as we can tell, does the post. That is not an accusation of bad faith. It is what a category looks like before anybody has run a controlled test, which is where this one currently is, ours included.
What is actually on page one for this term?
The live top ten for this head term contains nine explainer guides, seven of which are vendor or agency blogs selling the service they explain, one analyst house, and at position one a forum thread where somebody is asking the question rather than answering it. That composition is the argument of this post, and it is cheap for anyone to check.
We pulled the live top ten for the head term through a web search API on 2026-09-01, from a United States location, and read every result. The composition is the argument.
What is actually on page one for this term
| Position | Domain | Type | Source |
|---|---|---|---|
| 1 | reddit.com | Forum thread. A buyer asking peers, not a published guide. | measured |
| 2 | marketermilk.com | Vendor or agency blog, explainer guide. | measured |
| 3 | llmrefs.com | Vendor blog, explainer guide. | measured |
| 4 | idc.com | Analyst house. The only non-vendor authority in the ten. | measured |
| 5 | neilpatel.com | Agency blog, explainer guide. | measured |
| 6 | linkedin.com | Practitioner explainer on a professional network. | measured |
| 7 | vezadigital.com | Agency blog, comparison of the four acronyms. | measured |
| 8 | flow-agency.com | Agency blog, best-practices guide. | measured |
| 9 | resultfirst.com | Agency blog, comparison against traditional SEO. | measured |
| 10 | dwao.in | Agency blog, explainer guide. | measured |
n = 10 · as of 2026-09-01
Method: Top ten organic results pulled through firecrawl web search, location United States, for the exact head term, on 2026-09-01. Type is our reading of each page, not a field the search API returns. What would falsify it is simple and cheap. Run the same query yourself and count. If the forum thread is no longer at position one, or fewer than seven of the ten are vendor or agency blogs, the gap this post is built on has closed and the argument weakens with it.
Nine of the ten are explainer guides. Seven of the ten are vendor or agency blogs selling what they explain. One is an analyst house framing the subject as a shift rather than a tactic. And position one, above all of them, is a forum thread.
A forum thread outranking an analyst house and eight agency blogs on a commercial term is a specific signal, and it is not a signal about Reddit's domain authority. It is what a saturated definitional slot looks like when the published answers are not answering the question the searcher has. Google has ten candidate explanations of what LLM SEO is and it is choosing to put a conversation at the top.
Read the thread and the reason is not subtle. The advice in it is fine. Build a prompt set. Answer those prompts in as many places as you can. Make your content chunkable. Do digital PR. Get your leadership out in public. It is reasonable, it is broadly the same advice as the nine guides, and the author of the thread reads it, thanks everyone, and then asks this:
are you using anything to keep track of visibility? I know there are some paid services out there but idk which ones are worth it?
He goes on to name four products he has heard about from people on social platforms. Nobody in the thread answers that question with anything a buyer could act on.
Every guide for this term tells you what to do. Not one of them tells you how you would know it worked. That is not an oversight in the writing, it is the hardest part of the problem.
That is the gap. Nine guides answering the question of what to do, one thread asking the question of how you would know, and no published answer to the second. The rest of this post is the second question.
There is one thing we did NOT measure and will not pretend to. We did not probe whether an AI Overview was present on this result page, because doing that properly requires a specific probe configuration and without it the call returns zero references on queries that carry many. So the answer to "does this term trigger an AI Overview" is unknown here rather than no. An unmeasured thing reported as an absence is how most of the numbers in this category get made.
Rate limiting for production APIs
competitor.example › blog
roundup.example › guides
What the nine ranking guides get right
It would be easy to read the sections above as a dismissal of every other result on this page, and that would be both unfair and unhelpful. The guides are mostly correct about the tactics. They are missing one thing, and it happens to be the thing their reader asks for next.
So here is what to take from them, stated as generously as we can.
Build a prompt set before you build anything else. The highest-voted reply in the position-one thread, from a disclosed founder of a visibility tool, opens with exactly this:
Build a list of 20-50 prompts your target customers might ask.
That is the right first move and it is right for a reason the reply does not spell out. The prompt set is not a research artifact, it is the unit you will later split into two arms. A team that skips it has nothing to measure against ninety days later, whatever else it does. Twenty to fifty is also a sensible range for discovery, though note it sits below the forty-by-forty design the arithmetic above asks for once you move from looking to measuring.
Answer those prompts in more than one place. The same reply says do not restrict this to blog posts, and to put answers on community platforms, video, professional networks and question sites as well. Our own citation counts say the same thing from the evidence side: 91.5 percent of what we measured pointed off site. The guides reached the right conclusion by intuition and it happens to survive measurement.
Write chunkable, self-contained sections. Correct, and correct for humans too.
Do digital PR rather than link acquisition. Also correct, and it is the piece most teams are least equipped to do, which is why it stays undervalued.
Watch what the assistant searches for, not only what your buyer types. One commenter suggests asking the model to show its own decision process and the semantic questions it used, which is a rough but genuinely useful way to see the fan-out step described earlier in this post.
None of that is wrong. What none of it supplies is the design that would tell you whether doing it worked, which is the question the person who started the thread asked, and which is why a forum post is outranking every one of the guides that answers everything except it.
There is one specific thing to be careful with. Several of the guides recommend structuring content for machines: fragmenting pages into short standalone chunks, front-loading entity names, writing headings as literal prompt strings. Google's published guidance warns against this by name, saying there is no need to write awkward keyword-stuffed copy or chop content into tiny artificial snippets because the systems understand language the way a person does. When the platform's own advice and a vendor's advice diverge on a point the platform is in a position to know, the platform is the better bet, and the vendor version costs you a page that reads badly to the humans who are still most of your audience.
How does an assistant assemble an answer, and where can you touch it?
An assistant answer is built in five stages. The buyer's question is fanned out into several rewritten searches, those searches hit an index, candidate documents come back, the model writes prose from them, and separately it decides which of those candidates to name. You can act on four of those five, and the fifth, selection, is where the category's confident advice runs out of evidence.
Before the tactics, the mechanism, because almost every bad tactic in this category comes from a wrong mental model of what happens between a buyer's question and a named citation.
A buyer types a question. The assistant rarely searches for that exact string. It fans the question out into several searches, often rewriting it into the phrasings it expects to find documents under. Those searches hit an index. Candidate documents come back. The model then writes an answer from what it has, and separately decides which of those candidates to name.
Mechanism
How an answer gets assembled, and the four places you can touch it
- Buyer questionAn unbranded shortlist or comparison question, typed before they know you exist.
- Query fan-outThe assistant rewrites one question into several searches against the index.
- RetrievalCandidate documents come back. Most of them are not yours.
- Answer assemblyThe model writes prose and decides which of the retrieved sources to name.
- Named citationThe only event a reader sees and acts on.
- Buyer questionQuery fan-outrewritten
- Query fan-outRetrievalsearched
- RetrievalAnswer assemblycandidates
- Answer assemblyNamed citationselected
- RetrievalNamed citationretrieved but never named
Four of those five stages are places you can act, and they are not equally accessible.
The fan-out is where a lot of the useful, unglamorous work lives. If your page answers the question a buyer types but does not answer the rewritten version the assistant actually searches for, you are invisible for a reason that has nothing to do with quality. This is old keyword research wearing a new hat, and it is the single most transferable skill from classic SEO into this work.
Retrieval is ordinary search. Your page has to exist in the index, be crawlable, be about the thing, and be a plausible result for the rewritten query. Nothing here is new, which is exactly Google's point.
Assembly is where you have the least direct control and the most indirect. The model writes prose, and prose has a shape: it wants a claim, a reason and an example. A page that supplies a clean claim with a number attached and a source beside it is easier to lift into that shape than a page that buries the same fact in a paragraph of positioning.
Selection, the step where a retrieved source becomes a named one, is the part nobody can give you a reliable tactic for. It is also where the largest measured gap in this whole subject sits.
From the field
Retrieved and cited are different events, and somebody measured the gap
Across 1.4 million ChatGPT prompts, Ahrefs separated the URLs a model retrieved from the URLs it actually named in an answer. The two sets are very far apart. This matters for anyone buying LLM SEO, because a tactic can move retrieval for weeks with no visible change in citations, which looks like nothing happening right up until it looks like a step change.
That distinction between retrieved and named is worth holding onto, because it explains a pattern that otherwise looks like failure. A team does the work, the pages start getting pulled into context, and the citation count does not move for weeks. Nothing visible is happening. Then it moves in a step. If your instrument sees only named citations, and most do, that entire build-up is invisible to you, and the temptation is to abandon the work halfway through the period where it was actually taking effect.
There is a fifth place people try to act, which is the model itself, and it is worth naming so you can recognise the pitch. Training data is fixed at training time. You cannot get a page into a finished model, you cannot un-train an unflattering one, and any offer to do either is either confusion or a lie. What you can do is change what the retrievable web says about you between now and the next training run, which is a slower and less glamorous version of the same ambition.
The work, part one: the pages you own
On your own pages the work is ordinary craft done to an unusually high standard. Answer the question near the top in one self-contained paragraph, write sections that stand up on their own, put the number in the same sentence as the claim, and name the source at the point of use. This is the smallest of the three layers by citation share and the only one you control outright.
This is the layer everybody starts with, it is the smallest of the three by citation share, and it is still worth doing properly because it is the only layer you control completely.
Answer the question, near the top, in one paragraph. If the page is about a question a buyer asks, the first hundred words should contain a self-contained answer to it, phrased so it survives being lifted out of the page with no surrounding context. This is the highest-return single change on most pages and it costs an hour. Every page on this site carries one, which is the only reason we are comfortable prescribing it.
Make sections self-contained. The Reddit thread's own top advice puts it well:
Create content pillar pages with lots of cross linking to related information. Make sure that it is chunkable - ie each subtitle and paragraph are self contained.
That is correct, and it is also just good editing. A section that only makes sense after the three before it is bad for a skimming human as well as for an extractor. Note that this is not the same as chopping your page into artificial fragments, which Google warns against by name. The instruction is to write sections that stand up, not to write stubs.
Put the number in the sentence with the claim. "Most citations point off site" is an assertion. "Ninety-one and a half percent of the 901 citations we counted across 60 answers pointed off site, across 82 distinct hosts" is a quotable unit. The second one gets reproduced and the first one gets paraphrased away, and reproduction is the entire game.
Say where the number came from, at the point of use. This is the part almost nobody does and it is cheap. A figure with a source beside it survives being lifted into an answer with its attribution attached. A figure without one arrives in the answer as an orphan and is easy for the model to replace with somebody else's.
Structure and markup, briefly, because it gets oversold. Use real headings in a sensible order. Use schema where the page genuinely is the thing the schema describes. Google's guidance says explicitly that no special markup is required for AI features, so treat schema as good hygiene rather than as a lever, and be suspicious of anyone pricing it as one.
Depth, honestly assessed. The pages that get cited on competitive questions tend to be the ones that actually answer the question rather than positioning around it. That is not a word count target. It is that a page which handles the objection, states the limit and shows the working has more citable surface area than one that does not.
Freshness, with a caveat. Cited pages skew recent, and we have written about half of AI citations are under 13 weeks old elsewhere. The caveat is that the published evidence for refresh-as-tactic is correlational: pages got updated and citation counts moved, with no control group left deliberately unrefreshed over the same window. So refresh pages that are genuinely stale, and do not buy a refresh programme on the strength of that correlation alone.
Page one ranking
What the model read
Rank 1 appears in 0 of the 3 pages read
The work, part two: the pages you do not own
Most citations do not point at you. In our own count of 901 citations across 60 assistant answers, 91.5 percent pointed off site across 82 distinct hosts, which means the pages deciding your evaluations are roundups, review platforms, community threads and documentation you do not maintain. This layer is slower, harder to staff and where the return actually is.
This is where most of the citations actually are, and it is the layer that makes this work feel different from classic SEO even though the underlying mechanism is the same.
Where the citations we counted actually pointed
| Measure | Value | Source |
|---|---|---|
| Citations counted | 901, across 60 assistant answers | measured |
| Share pointing off site | 91.5 percent | measured |
| Distinct hosts cited | 82 | measured |
| Share pointing at the vendor's own domain | 8.5 percent | derived |
n = 901
Method: Counted from our own live-answer measurement runs on our own money queries. A citation counts when a source is named in the answer text a reader sees, not when it merely appears in a collapsed source panel. The last row is arithmetic on the third, which is why it is tagged derived rather than measured. This would be falsified by a query set of the same shape returning a materially lower off-site share, and the query set is the thing to argue with first.
Ninety-one and a half percent. Across 82 distinct hosts. On our own money queries, using our own measurement, with a deliberately narrow definition of what counts as a citation. If we used a looser definition the number would be different, which is precisely why the definition is written down.
The consequence is uncomfortable for anyone whose plan is a content calendar. You can publish excellent pages forever and still lose the questions that matter, because the pages deciding those questions belong to review sites, comparison roundups, forums, documentation you do not maintain and issue threads where somebody else described your product's failure modes in public.
So the work splits into surfaces, and they are not equally tractable.
Comparison and alternatives pages, on third-party domains, are the highest-value and hardest target. When a buyer asks for the best tool for a job, the assistant frequently reads a roundup. Getting into a roundup you do not own is outreach, and it is slow, and there is no shortcut that is not transparently a shortcut.
Review platforms are the most mechanical. A category page with your product missing is a page that cannot cite you. This is unglamorous list-maintenance work and it is frequently the fastest thing on this list to fix.
Community threads are the largest and the most dangerous. Reddit's own published figures put it at around 21 percent of Google AI Overview citations and 46.7 percent of Perplexity responses, which is a real share, and it is why an entire service category has grown up around posting there.
From the field
Community pages are overrepresented, on the platform's own disclosure
Reddit's own published figures put it at around 21 percent of Google AI Overview citations and 46.7 percent of Perplexity responses. That is a real and large share, and it is why so much of this category has become Reddit advice. What no published version of that advice includes is a query set, a control group or a sample size, so the step from running a campaign to being recommended is untested rather than proven.
Reddit, how we are keeping Reddit real and safe in the AI era
We have taken that claim apart at length in what the Reddit citation data does not prove, and the short version is that the platform share is well evidenced and the step from "we ran a Reddit campaign" to "we got recommended" is not. If you do this work, and there are good reasons to, the posting rules that keep it survivable matter more than the volume.
Documentation, quickstarts and issue threads, for anyone selling to developers. A model reads an API reference as an answer rather than as marketing, and it reads a competitor's issue tracker as evidence about that competitor. If you sell a developer tool or an API, the surfaces that decide your evaluations are mostly technical and mostly not yours.
The thread we have been quoting picked up the same point from the other end, and it is a fair statement of what this layer costs:
Offsite comments will be huge going forward. We were just discussing the pivot to digital PR.
Digital PR is the honest name for most of this. It is slower than publishing, it is harder to staff, and it is where the citations are.
82 distinct hosts carried them between them
No single site owns a category, so there is nothing to buy your way onto. Our own measurement.
The work, part three: what other people say about you
Underneath the pages you own and the pages you do not sits the evidence an assistant assembles from. There is no stored record of your company waiting to be read; a version of you is built at the moment of asking out of whatever gets retrieved. That makes public conduct, named authorship and earned coverage an input to this work rather than a nice-to-have beside it.
There is a third layer that sits underneath both of the others, and it is the least tactical thing in this post.
An assistant does not hold a record of your company. It assembles one, at the moment of asking, out of whatever it retrieves. A public relations practitioner put this well while citing Semrush's 2026 AI Visibility Index, and we quote it as reported because we have not read the index ourselves:
LLMs build a new version of you every time someone asks.
The figures he cites from that index are worth repeating with the same caveat attached. It covers 126 million AI search prompts across four surfaces and tracks more than 1,200 brands, and only 36 of them stayed among the 100 most-mentioned on every platform in every month studied. All 36 were household names. He also reports that ChatGPT cited an average of 15 sources per response against three for Gemini.
Take the divergence figure seriously even at second hand, because it lines up with our own citation counts: the surfaces disagree, and a blended number across them describes a market nobody sells into.
The practical form of this layer is not a tactic, it is a posture, and the Reddit thread put it plainly:
Honestly, your leadership, product managers, SMEs are going to have to get out and promote themselves and the company
That is unfashionable advice and it is correct. Podcasts, conference talks, public writing under a real name, answering questions in the places your buyers already are. It generates the retrievable evidence that the layer above is made of. It also cannot be bought as a deliverable, which is why almost no vendor page in this category leads with it.
You cannot edit what a model was trained on. You can change what the pages it retrieves today say about you, and those are two different projects with two different timelines.
Branded-win, generic-invisible
What is outside your control, and saying so
Four things in this system are genuinely beyond reach. Training data is fixed at training time, retrieval is stochastic, the selection step is opaque from outside, and provider model updates land whenever they land. Naming those honestly is not a disclaimer, it is what separates a method you can hold somebody to from a promise nobody can check.
The most useful comment in the position-one thread is the one that admits a limit. It is worth quoting in full because it is more honest than most published guides on this term:
The harsh reality is this is partially outside your control since you can't directly influence what LLMs were trained on. But you can influence what authoritative sources say about you going forward, which may impact future model updates and retrieval-augmented systems.
Both halves are right, and the second half is the whole business. But the first half deserves more weight than the category gives it, so here is the honest inventory of what you cannot touch.
Training data is fixed. Whatever the web said about you at the cutoff is in there, including the things you would rather it did not say, and no amount of publishing changes it. New pages affect retrieval now and may affect a future model. Those are different timelines and it is worth being clear with a stakeholder about which one you are promising.
Retrieval is stochastic. Ranking is not deterministic, the index is being rewritten continuously, and the same question asked twice does not necessarily reach the same documents.
Selection is opaque. Even holding retrieval constant, which source gets named is a decision inside the model and nobody outside the lab has a reliable account of it.
Model updates land whenever the provider ships them, and they can move your numbers by more than your work did, in either direction, with no announcement.
And there is a structural point that one commenter in the thread made and that most vendor writing avoids entirely:
it does beg the question of how you drive traffic in a world where LLMs gobble content and don't send traffic or play to dominant players in any give space.
That is a real question and we do not have a complete answer to it. Being named in an answer is not the same as being visited, and a citation strategy that cannot connect to a pipeline eventually has to justify itself on something other than a citation count. We think the connection exists on shortlist questions specifically, because being one of three named options at the moment somebody is choosing is worth something even without a click. We cannot currently prove it with our own data, and we would rather say that than assert it.
Overview fired
5 of 5
Median sources
5
New brand cited
0 of 5
A representative five-trial pattern for a domain with no publication history yet, the Overview fires nearly every time and a newcomer is still named in none of them, until authority accumulates elsewhere.
The question the top-ranked thread actually asks
After reading a page of good tactical advice, the author of the position-one thread asks which visibility tracker is worth paying for. That is a measurement question wearing a shopping question's clothes, and it is the right question at that point. Choosing a tool is the wrong proxy for it, because no tool on its own can tell you whether your work caused a change.
Everything above is the tactical half, and it is the half nine other pages already cover. Here is where the post turns, because the thread at position one turns here too.
The author reads the advice, thanks the people who gave it, and asks which tracker is worth paying for. He names four products he has heard about. He does not get an answer.
That question is not a shopping question. It is a measurement question wearing a shopping question's clothes, and it is the correct question to ask at that point in the process. He has a list of things to do. He wants to know how he would tell whether doing them worked. Choosing a tool is his proxy for answering that, and it is the wrong proxy, because no tool answers it on its own.
Here is why. Nearly every tool in this category reports a level. A visibility score, a share of voice, a mention count, read at a point in time. That is a perfectly reasonable thing to build and it is a genuinely useful diagnostic. It is not a measurement of your work, and the reason has nothing to do with the quality of the tool.
Two things have to be true before a level means anything. The reading has to be stable enough that observing it once tells you where you are. And there has to be something to compare it against that would have moved the same way without you.
Neither is true here. The next two sections are the evidence for that, and they are the two sections the other nine results on this page do not have.
What a real read looks like
A two-arm citation read, at the point where it becomes readable
40
Queries, treated arm
40
Queries, held out
40
Samples per query
9.9 pts
Minimum detectable lift
- Query set chosen by the buyer, unbranded, published in full
- Arms split and written down before any baseline is taken
- Per-engine breakdown, never a blend
- Citation rule stated in writing before counting
- A causal lift numberWe do not have one yet. The instrument passed its kill test; the 60-day test has not run.
Your buyer asks
The answer they get
For production workloads, most teams land on Competitor API1. It pairs token-bucket limits with per-key analytics.2
Where the citations resolve
Why does one reading of a visibility score behave like a coin toss?
An assistant answer is not stable between identical asks. We ran twelve money queries five times each on one model in one day with nothing changed between runs, and seven of the twelve changed outcome. A single reading of where you stand is therefore not a weak measurement, it is a draw from a distribution with a decimal point printed on it.
We tested this on ourselves before building anything on top of it, because the whole product depends on the answer.
Twelve money queries. Five identical repeats each. One model, one day, nothing changed between runs. Sixty calls, all sixty succeeded.
Seven of the twelve queries changed their outcome across those identical repeats.
Not changed wording. Changed outcome: named in one run, absent in the next, same question asked the same way minutes apart. If a vendor had run one of those queries once and shown you the result, you would have had close to even odds of seeing the opposite finding.
This is not a defect in a particular model and it is not a vendor cutting corners. Retrieval is stochastic, the index moves, and the answer is generated rather than looked up. Instability is the system working as designed. Published work shows that even temperature zero is not deterministic, with accuracy varying by up to 15 percent across ten runs on the same task, so this is not a setting somebody forgot to switch off.
Nor is it only us seeing it. SparkToro, testing repeated identical prompts, put the chance of getting the same brand list twice at under one in a hundred.
Now add the churn underneath it.
Worth knowing
The source list churns while the answer stays the same
Ahrefs, measuring across 43,000 keywords, reports that 45.5 percent of citations change between consecutive observations while the answer itself stays about 95 percent semantically identical, and that AI Overviews persist for around 2.15 days. The wording holds steady and the sources underneath it move. A before-and-after comparison run against a set that reshuffles this fast will show movement whether or not anybody did anything. We carry this figure with its attribution rather than a link, because our own source registry does not record a public URL for it and guessing one would be worse than saying so.
Ahrefs, citation churn across 43,000 keywords
Put those two facts together and the standard case study in this category stops working. Measure a score. Do some work. Measure again. Report the difference. If a single reading can flip on its own, and the source list underneath the answer reshuffles between observations anyway, then the difference between two readings contains the work you did plus however much the system moved by itself. Nothing in that method separates the two.
A vendor showing you a twenty-point improvement has not shown you that they caused twenty points. They have shown you two draws from a distribution.
The uncomfortable version: run the same before-and-after with no work done at all and you will still get a number. Sometimes a flattering one. We have written the long form of this argument up separately in why a single AI visibility score is noise, with the full run and the arithmetic.
This is the point where the category mostly concluded that AI citations cannot be measured and stopped. That inference is wrong, and the reason it is wrong is the next section.
best rate limiting api
The two-arm design, and what it costs to run
The fix for an unstable metric is not more precision on one reading, it is a comparison. Split your query set, work one half, deliberately leave the other half alone, measure both at the same time, and report the difference between them. Whatever the system did on its own it did to both arms, so the shared movement subtracts out and what remains is attributable.
The noise in the system is unbiased. That is the load-bearing claim, so here is exactly what is established and what is assumed.
That the noise exists is measured, by us and by others. That it is symmetric and independent across a treated and a held-out arm of AI-visibility queries is the standard assumption of controlled experiment design, and we have not found anybody who has tested it directly on citation data. We assume it, we tested it on our own query set, and we are telling you it is an assumption rather than a finding.
We tested it rather than asserting it. Across 20,000 random splits of our own query set with no intervention applied at all, the difference between the two halves centred on zero: mean plus or minus 0.0016, standard deviation 0.215. Under the null hypothesis, the difference between two arms is genuinely zero even though each arm individually is jumping around.
That is the whole trick. You cannot trust a level. You can trust a difference between two groups measured the same way at the same time, because whatever the system did to one arm it did to the other.
The design itself is one decision made before you start. Take the query set you care about and split it. Work one half. Leave the other half alone, deliberately, for the whole period. Measure both at the same time, with the same prompts, on the same day, at the same trial count. Report the difference between the halves, never the change in the worked half.
A number that only ever goes up is not a measurement, it is a report on a system nobody is holding still. The held-out arm is the whole of the difference between the two.
Once you accept two arms, sample size stops being a detail and becomes the constraint.
What each measurement design can actually detect
| Design | Calls per timepoint | Minimum detectable lift | Verdict | Source |
|---|---|---|---|---|
| 12 queries by 5 samples | 60 | 51.1 points | Our own pilot. Cannot see a real engagement. | derived |
| 20 queries by 20 samples | 400 | 21.4 points | Still coarser than the effect being measured. | derived |
| 40 queries by 40 samples | 3,200 | 9.9 points | Usable. Roughly six hours of machine time, twice. | derived |
Method: Computed from the variance in our own step-zero run rather than taken from a vendor table. A real engagement moves the number by five to fifteen points, which is why the first row would report almost every genuine win as nothing. Falsified if the variance on a different query set is materially lower, in which case the required sample sizes fall and the first row becomes defensible.
Statistical power
Minimum detectable lift by design size
| Point | Value (percentage points) |
|---|---|
| 12 queries by 5 samples | 51.1 percentage points |
| 20 queries by 20 samples | 21.4 percentage points |
| 40 queries by 40 samples | 9.9 percentage points |
Our own twelve by five design has a minimum detectable lift of 51.1 percentage points. That is not a measurement instrument. A real engagement moves a citation share by something like five to fifteen points, so a design that can see only a fifty-one point move reports almost every genuine win as nothing at all.
Getting the floor to 9.9 points takes forty queries by forty samples, which is roughly 3,200 calls per timepoint and about six hours of machine time. Twice, because you need a before and an after.
That cost is why almost nobody runs it. It is also why a number produced this way means something. If you want the arithmetic in full, including how the minimum detectable lift is computed, it is written up on how we measure lift.
One condition, easy to state and easy to forget: the two arms have to have been moving in parallel anyway. A model update mid-window that lands harder on one arm's queries than the other's breaks that, and no amount of sampling repairs it. Per-call sampling variance cancels in a difference. A systematic behavioural change correlated with time does not, if it hits the arms unevenly.
And the honest disclosure that belongs here rather than at the bottom:
What to ask before you pay for a tracker
This is the direct answer to the question at the top of the search results, and it is a set of questions rather than a product name, because a product name would be worth less to you.
Five questions to put to any LLM SEO tool or agency
| Ask | What a real answer contains | What silence means | Source |
|---|---|---|---|
| Which query set | The full list, chosen by you, unbranded, questions a buyer types before they know you exist. | The number describes their market, not yours. | unknown |
| Which engine, per figure | A per-surface breakdown, never a blend. | The figure mixes surfaces and only one of them is in your pipeline. | unknown |
| How many trials per query | A stated repeat count, and the variance across repeats. | A single probe cannot tell a change from a repeat. | unknown |
| What counts as a citation | A written rule covering named in text, listed in a source panel, and retrieved but not named. | Two reports using different units are not comparable and neither is wrong. | unknown |
| Where is the held-out arm | A set of comparable queries deliberately left alone over the same window. | Every figure is a level. There is no answer to compared to what. | unknown |
Method: This table is prescriptive rather than measured, and every row is tagged unknown for that reason. It is the checklist we run on ourselves, not a survey of what vendors answer. We have not surveyed the category, so we do not know what share of tools would pass, and inventing a number for that would be exactly the behaviour the table exists to warn against.
Five questions. None of them requires statistical training to ask. A vendor who answers all five is doing real work whoever they are, and a vendor who cannot answer the last one is selling you a level and calling it a lift.
The reason we will not simply name a product is that the answer depends on which of three shapes you are buying, and they fail in different ways.
01 / Do it in house
- Stands out
- You own the query set, so nobody can pick the questions that flatter them. Cheapest by a wide margin if you already have somebody who can run scripted API calls.
- Best for
- Teams with an engineer who can spare a day a month, and a marketer who will not quietly drop the held-out arm when the quarter looks bad.
- Falls short
- The held-out arm is the first thing to get abandoned under internal pressure, and there is nobody outside the room to notice. Also roughly 3,200 calls per timepoint in metered API spend that lands on somebody's card.
02 / Buy a visibility tool
- Stands out
- Fastest to a number, and a good tool covers more engines and more queries than you would bother to script yourself.
- Best for
- Teams who need a recurring read across several surfaces and have somebody who will actually interrogate the methodology rather than screenshot the score.
- Falls short
- Most publish a single blended score, and a score you cannot decompose by engine and by query cannot tell you whether a move was you or the model. Ask the five questions in the table above before the trial ends, not after.
03 / Hire an agency
- Stands out
- The only option that also does the off-site work, which is where most of the citations actually live and where an in-house team usually runs out of road.
- Best for
- Teams who need execution across surfaces they do not own and cannot staff for it.
- Falls short
- The party doing the work is also reporting on it, which is the oldest conflict in marketing and does not get better because the surface is new. Make the query set and the held-out arm contractual, or the report is a level rather than a lift.
Two notes on reading that honestly. We are the third shape, so treat the weakness we wrote against ourselves as the most load-bearing line in the three. And we have written a longer comparison of the best AI visibility tools for 2026, measured, plus a head-to-head against the category's best-known vendor, both of which are more useful than a ranking would be here.
The one thing worth saying flatly: a tool and a measurement are different purchases. A tool tells you what the answers currently say, which is genuinely useful and is the right first spend. A measurement tells you whether your work changed them, and that needs a held-out arm, which no tool can create for you because the split is a decision about your own query set.
When this work is not worth doing yet
A post like this has an obvious incentive to tell every reader that they need the thing it describes. Here is the opposite, which is more useful and which we have not found on any of the nine guides.
Skip this entirely if nobody asks an assistant about your category. Not every purchase runs through one. If your buyers find you through a sales team, a partner channel, a physical location or a procurement list, then citation share is a vanity metric with a research budget attached. The cheap way to check is to ask ten recent customers how they first heard of you, before you spend anything.
Skip it if your classic search presence is broken. Google says its AI features retrieve from the existing index and run on the same ranking and quality systems. A site that is not crawlable, not indexed, or ranking nowhere for its own category terms does not have an LLM SEO problem, it has an SEO problem, and the AI layer sits on top of the thing that is not working. Fixing the foundation is also the cheaper project, which is the rare case of the boring answer being the profitable one.
Skip the measurement, but not the work, if your query set is small. The arithmetic is unsentimental here. If your category genuinely only supports twelve distinguishable buyer questions, a two-arm design over that set has a minimum detectable lift around fifty-one points, and no amount of extra sampling on twelve queries fixes it. Do the work, because the work is good practice regardless, and be honest that you are doing it on judgement rather than on measurement. That is a defensible position. Pretending the number you read means something is not.
Wait if you cannot protect a held-out arm. This is an organisational condition rather than a technical one. If the person paying for the work will not tolerate a set of questions being deliberately left alone for a quarter, you will not get a readable result, and you will have spent the money to produce a level. Better to know that in week zero.
Be sceptical of urgency framing generally. The loudest version of this category's pitch is that your competitors are being recommended right now and every month of delay is compounding. Some of that is true and some of it is a shortlink. The measured facts we can offer are that the source lists churn between observations, that most of the citations are on pages you do not own, and that nobody has yet published a controlled result showing which tactic moves them. None of those support a panic.
Two accounts, one script, a quarter of a million views
Reading the market matters as much as reading the evidence. While pulling public discussion for this post we found two large accounts carrying byte-identical promotional copy about this term, three days apart, with different shortlinks. That is worth knowing about the volume on this subject, and it is worth being precise about what it does and does not prove.
A short section on reading the market, because a large part of doing this work well is not being moved by the loudest version of the argument.
While pulling market voice for this post we ran an advanced search on the head term and read the top twenty results. Two of them, from separate accounts with 156,537 and 76,874 followers, carried byte-identical promotional text three days apart, differing only in the shortlink.
Two accounts, one script
| Account | Posted | Views when read | Source |
|---|---|---|---|
| Account A, 76,874 followers | 2025-07-04 | 137,784 | measured |
| Account B, 156,537 followers | 2025-07-07 | 115,951 | measured |
| Combined across both posts | Three days apart | 253,735 | derived |
as of 2026-08-31
Method: Both posts read through the X API on 2026-08-31 by tweet id. The promotional body text is byte-identical between them and only the shortlink differs. Counts are the view counts the API returned at read time and will have moved since. This is an observation about the corpus and nothing more. It is not evidence of who wrote the script or who paid for it, and it must not be read that way.
The text reads: the future of SEO is already here, it is called LLM SEO or LEO, and it is quietly driving hundreds of thousands of users, with one team shipping it for months and numbers to prove it.
We want to be precise about what this is and is not evidence of, because the temptation to overread it is strong and the overreading would be exactly the sin this post is about.
It is evidence that identical promotional copy about this term reached at least 253,735 impressions across two accounts. That is measured, from the API, on the date stated.
It is not evidence that the accounts were paid, coordinated, or acting in bad faith. Identical text appears for many reasons, including an affiliate programme supplying copy, and we did not check which applies here. Writing it as coordination would be inventing a fact to make a better paragraph.
What it is useful for is calibration. When you read a confident claim about this category, ask whether the person making it has a query set, a control group and a sample size behind it, or a shortlink. Most of the volume on this term is the second thing. That is not unique to LLM SEO, it is what every new marketing category looks like early, and the defence is the same as it always was: ask for the method.
6 of 10 answers
A ninety-day procedure you can actually run
Everything above reduces to six steps, three of which happen before any work starts. Choose forty unbranded buyer questions, split them into two arms and write the split down, take a baseline across both, work one arm for ten weeks, read both again identically, and report the difference. The steps that get skipped are always the ones in week zero.
Everything above collapses into a procedure. It is deliberately boring and it is the shortest honest version we can write.
STEPS
A ninety-day procedure you can actually run
Week 0, choose the questions
Half a day
Write 40 unbranded questions a buyer types before they know you exist. Drop anything with your brand in it, anything that returns prose with no vendor names, and anything asked after the decision rather than during it.
Week 0, split the arms and write it down
One hour
Randomly split the 40 into a treated arm and a held-out arm. Record the split before you take a single baseline reading, because a split chosen after the fact is not a split.
Week 1, take the baseline
Roughly six hours of machine time
40 samples per query across both arms, per engine, never blended. Record the citation rule you are counting under, in writing, before the first call.
Weeks 2 to 11, do the work on one arm only
Ten weeks
Owned pages, off-site surfaces, and the entity layer. Leave the held-out arm completely alone, which is harder than it sounds and is the step people quietly skip.
Week 12, take the second reading
Roughly six hours of machine time
Identical prompts, identical trial count, identical day, both arms. Report the difference between the arms, never the change in the treated arm.
Week 12, publish the null if it is null
One hour
If the difference does not clear the minimum detectable lift, say so. A null tells you the money should go somewhere else, and it is the outcome nobody publishes.
Six steps. Three of them are decisions made before you do any work at all, which is the part teams skip, and skipping them is what makes the result unreadable ninety days later.
A few notes on the steps that go wrong most often.
The split gets abandoned. Somebody looks at the held-out arm in week six, sees it is not moving, and quietly starts working it. This is the single most common failure and it is unrecoverable: once both arms are treated there is no comparison left, and the ninety days produce a level like everybody else's.
The query set gets chosen to flatter. A query set full of questions you already win is a query set that cannot show a lift. A query set with your brand name in it is measuring existing awareness back to yourself. Both are easy to do by accident and both are hard to spot afterwards, which is why the list gets written down in week zero and not edited.
The citation rule drifts. Counting a source panel appearance in the baseline and only named mentions in the second reading produces a decline that is entirely definitional. Write the rule down first. Ours is narrow, named in the answer text a reader sees, and narrow rules produce smaller numbers, which is a reason to publish the rule rather than to loosen it.
The engines get blended. Averaging ChatGPT, Perplexity and Google's surfaces produces a number whose volatility belongs to none of them and which cannot tell you where you are losing. Keep them apart from the first reading.
And the null gets buried. If the difference between the arms does not clear your minimum detectable lift, that is a result. It says the money should go somewhere else, which is more valuable than most positive findings and is the outcome nobody in this category publishes.
If you want this run against your own money questions rather than described, that is the answer engine optimization engagement and the design above is exactly what it does. If you would rather run it yourself, everything needed is in this post and we would genuinely rather you did that than bought a level from anybody.
one money query, nothing changed between runs
7 of 12 queries moved outcome across identical repeats in our own step zero run, 60 of 60 calls successful. One read is not a reading.
What would prove this post wrong?
A post arguing for falsifiable claims owes you its own falsification conditions, written in advance rather than assembled afterwards to fit whatever happened. Here are the five findings that would damage this argument, ordered by how much damage each one would do, along with what we would have to change rather than defend in each case.
Written before the results section rather than after it, because a criterion invented to fit an outcome is not a criterion.
What would change our mind, written before the argument rather than after
| Finding | What it would do to this post | Source |
|---|---|---|
| Google states its AI features use a separate ranking stack | Breaks the central claim. The tactics would genuinely diverge and most of this post would need rewriting rather than defending. | unknown |
| A vendor publishes repeat data showing assistant answers are stable | Narrows it hard. Stability is a property of a query set rather than of the category, and if a real set is stable then a single reading starts carrying information again. | unknown |
| Noise turns out to be asymmetric across a treated and a held-out arm | Breaks the measurement half. The difference would stop cancelling and the design would need repairing. This is the one we would most like somebody to test. | unknown |
| Someone publishes a controlled result showing a specific tactic moves citations | Strengthens the post rather than refuting it. We would cite it, because there is currently no such result to cite, including ours. | unknown |
| A reported visibility number goes up after some work | Nothing at all. That outcome is compatible with the work, with drift, and with two draws from one distribution. | derived |
Method: Written before the results section rather than after it, because a criterion invented to fit an outcome is not a criterion. The last row is tagged derived because it follows directly from our own permutation result rather than from a hypothetical.
The first row is the one that matters most. This entire post rests on Google's own statement that its generative features run on the shared ranking and quality systems and retrieve from the existing index. If the platform reverses that, or somebody demonstrates a separate stack, then the tactical half of this post is wrong and most of it would need rewriting rather than defending.
The third row is the one we would most like somebody else to test, because we cannot test it against ourselves without marking our own homework. If the noise in assistant answers turns out to be asymmetric across a treated and a held-out arm, the difference stops cancelling and the design in this post needs repairing.
The last row is the one to notice, because it is the shape almost every case study in this category takes. A reported number going up proves nothing. It is compatible with the work, with drift, and with two draws from one distribution, and there is no reading of it that distinguishes the three.
What can we still not measure?
The honest closing inventory, because a post arguing for measurement should say where its own instrument stops. Five things sit outside what we can currently observe, and each of them limits a claim somebody in this category is making confidently today, ours included. Naming them is cheaper than being asked for them later.
We do not have a causal lift result. The instrument exists and passed its own kill test, the design is pre-registered, and the sixty-day two-arm test has not run. Any number we published today would be the thing this post spends nine sections arguing against.
We cannot measure the effect on pipeline. Being named in an answer is not the same as being visited, and we cannot currently connect a citation to a deal with anything better than a story. We think the connection is real on shortlist questions specifically. We cannot prove it, and neither, as far as we have found, can anybody else.
We cannot see inside selection. We can measure that a source was named and that another was not. Why the model chose one over the other is not visible from outside, and every published account of it we have read is inference rather than observation.
We cannot generalise a query set. Stability is a property of a particular set of questions rather than of the category. Ours flipped seven times out of twelve. Yours might be steadier or worse, and the only way to know is to run the repeats.
And we cannot measure the training-data layer at all. What a model absorbed before its cutoff is fixed and unobservable from here, which means part of what determines whether you get named is permanently outside both your control and your instrument.
That list is longer than we would like. It is also, as far as we can tell from reading the nine other results on this page, the only such list on any of them. We would rather publish it than compete on confidence.
If you are looking for the sequence rather than the definition, how to get an API recommended by AI assistants sets out the three moves in the order that pays. And if your instinct is to start with your reference docs, that is the one surface that cannot win a choosing question. One scope question sits above all of it: whether any of this applies outside developer tools and APIs, which is worth checking before assuming the rest of this post transfers to your category.
Sources
Every number above, and where it came from. A figure without a row here is one we should not have printed.
- Google, Brendon Kraham on generative AI features, reported by Glenn Gabe
- The dispositive first-party statement in this post. Google's own position is that its generative AI features are rooted in the same core ranking and quality systems as traditional Search, and retrieve up-to-date content from the existing search index. Read live via the X API on 2026-08-31, full text, not truncated.
- Google Search Central, optimizing your website for generative AI features
- Google's own published guidance, the corroborating document behind the statement above. States that generative AI features on Search are rooted in core Search ranking and quality systems and that no special markup is required.
- r/AskMarketing, effective SEO and GEO strategies for LLM visibility
- The position one organic result for the head term this post targets. Read via the Reddit API on 2026-08-31, 37 comments excluding the AutoModerator post. The thread author's own follow-up question is which visibility tracker is worth paying for, which no ranking guide answers.
- Jake Ward, KPIs from Google SEO will not help with LLM SEO
- The strongest opposing case to this post's argument, engaged rather than dismissed. Read live via the X API on 2026-08-31.
- Sarvesh Shrivastava, local SEO is being replaced by LLM SEO
- The replacement claim in its purest published form, used here as evidence of how the term is being sold rather than as evidence about search behaviour.
- Two accounts, one identical promotional script
- Byte-identical promotional text posted by two large accounts three days apart with different shortlinks, 115,951 and 137,784 views respectively when read on 2026-08-31. Recorded as an observation about the corpus. It is not evidence about who coordinated it, and is not written here as though it were.
- Bob Pickard, on what generative systems assemble about a brand
- A public relations practitioner citing Semrush's 2026 AI Visibility Index, 126 million prompts across four surfaces and more than 1,200 brands. Quoted as reported. We have not read the underlying index ourselves and say so at the point of use.
- Lawrence Hitches, LLM SEO, how to optimize for AI-driven search engines
- The professional-network result in the live top ten for this term, included as evidence of what currently ranks rather than as a source for any figure.
- Ahrefs, why ChatGPT cites one page over another
- 1.4 million ChatGPT prompts. Separates URLs a model RETRIEVED from URLs it CITED, which is the distinction most of this category collapses.
- Reddit, how we are keeping Reddit real and safe in the AI era
- Reddit's own disclosure, the primary source for the platform-share figures quoted here.
- SparkToro and Gumshoe, repeated-prompt consistency
- Under a one in a hundred chance that two identical prompts return the same brand list. An independent figure of the same kind as our own repeat data, on a different corpus.
- Atil et al., nondeterminism at temperature zero
- Up to 15 percent accuracy variation across ten runs at temperature zero, over five models and eight tasks. This is why the instability is not a setting somebody forgot to turn off.
- Our own step-zero measurement run
- 12 money queries by 5 identical repeats, one model, one day, 60 of 60 calls succeeded, 7 of the 12 flipped outcome across identical repeats. First-party, no public URL.
- Our own permutation test and power analysis
- 20,000 random splits of the same query set with no intervention applied, null difference centred on zero, mean plus or minus 0.0016, standard deviation 0.215. Minimum detectable lift of 51.1 points for the 12 by 5 design, 21.4 for 20 by 20, 9.9 for 40 by 40. First-party.
- Our own citation-destination count
- 901 citations across 60 answers, 91.5 percent pointing off site, across 82 distinct hosts. First-party, and the figure this site uses in place of the unsourceable industry version.
- Our own live SERP pull for this term
- Top ten organic results for the head term, pulled 2026-09-01 through firecrawl web search, United States. Nine explainer guides, seven of them vendor or agency blogs, one forum thread at position one, one analyst house. First-party.
- Ahrefs, citation churn across 43,000 keywords
- 45.5 percent of citations change between consecutive observations while the answer itself stays about 95 percent semantically identical, and AI Overviews persist 2.15 days. Carried here with its attribution because our own source registry does not record a public URL for it, which is stated rather than papered over with a guessed link.
Questions people actually ask about LLM SEO
- What is LLM SEO?
- Getting a brand named and cited inside AI assistant answers rather than only ranked in a list of links. Google says its generative AI features run on the same core ranking and quality systems as Search and retrieve from the existing index, so most of the on-page work is ordinary SEO done properly.
- Is LLM SEO different from GEO and AEO?
- Not meaningfully. GEO frames the same work around generative surfaces, AEO is older and covers snippets and voice answers too, and LEO is mostly a coinage attached to a course. Google's own position is that the new names describe good SEO. Pick one and be consistent.
- Do I need new tactics for LLM SEO?
- Mostly no. Clear structure, self-contained sections, a real answer near the top and genuine off-site presence were already good practice. Google warns explicitly against awkward keyword-stuffed copy and artificially chopped snippets, so the machine-shaped version of this advice is worse than the ordinary version.
- How do I measure whether LLM SEO worked?
- Split your query set into a treated arm and a held-out arm, write the split down before the baseline, sample every query many times per engine, and report the difference between the arms rather than the change in the treated one. A single reading cannot separate your work from the system's own movement.
- Why does my AI visibility score keep moving on its own?
- Because assistant answers are not stable between identical asks. In our own run, 12 queries repeated 5 times each on one model in one day produced 7 that changed outcome with nothing altered between runs. Independent work puts repeated-prompt agreement under one in a hundred.
- Can I influence what a model was trained on?
- No. Training data is fixed at training time and outside your control. What you can influence is what the pages a model retrieves today say about you, which affects retrieval-augmented answers now and may affect later model updates. Those are two projects with two timelines.
- How many samples does an LLM SEO measurement need?
- It depends on the effect you expect. Our own 12 by 5 pilot detects only a 51.1 point lift, which is useless. Forty queries by forty samples brings the floor to 9.9 points at roughly 3,200 calls per timepoint. There is no single sample count that works for every query set.
Keep reading
AI citations
Reddit is cited by AI. That is not a reason to buy upvotes.
Published Reddit citation shares run from 2% to 46.7%, and one study puts Reddit at 67.8% of every URL ChatGPT retrieves and then declines to cite.
23 min read
AI citations
Does llms.txt actually move AI citations, and how would you tell
Every page ranking for this term explains what llms.txt is. None answers whether it changes anything. What 40 sites deploy, and the test that would tell.
43 min read
AI citations
What Predicts an AI Citation Is Not What Proves Yours
DiscoveredLabs measured 2 million AI citations and found what correlates with getting cited. Correlation across pages you don't own isn't proof for yours.
22 min read