Schema Markup for AI Search: How to Isolate the Variable (2026)
Ahrefs tracked 1,885 pages adding schema and citations barely moved. Here is how to isolate structured data as a variable, and what a null result looks like.
Citon80 min read

Short answer
Does schema markup cause AI citations?
Google states there is no special schema.org structured data you need to add to appear in AI Overviews or AI Mode. The largest controlled test, 1,885 pages against 4,000 matched control pages, found no citation uplift on any platform. Schema's effect on AI citations, if any, is indirect, and no public test has isolated it.
On 2026-08-27 somebody asked a question in r/TechSEO that page one of Google has never answered. The thread was small, 14 upvotes, and it drew 28 comments, which is the ratio you get when a question is contested rather than settled. The sentence in it that matters was six words long.
Has anyone here actually isolated schema as a variable.
Nobody had. Ten pages rank for this query in the United States and we read all ten on 2026-09-09. Nine of them carry no original measurement at all. Two carry causality in the title and run no comparison in the body. The one page that did run an intervention reports a result that most of the pages citing it get backwards.
This post answers that question directly. Not "does schema work", which is unanswerable as asked, but the narrower and more useful version: what would you actually have to do to find out, and what would the answer have to look like before it meant anything. We have run the measurement half of that on our own query set and we publish the numbers below, including the number that says our own instrument is not yet good enough.
At a glance
What the public evidence on schema and AI citations actually is
| Study | Design | Control group | What it found |
|---|---|---|---|
| Ahrefs, May 2026 | 1,885 treated pages, difference-in-differences | Yes, 4,000 matched pages | No uplift on any platform |
| searchVIU, Dec 2025 | 8 extraction tests, 5 AI systems | Yes, a visible-HTML baseline | JSON-LD-only fact found by 0 of 5 |
| Otterly, Mar 2026 | 319 prompts, 7 platforms, one site | Informal, competitor parallel move | Effect indirect, 6 of 7 could not read it |
| Kumar and Palkhouski, Sep 2025 | 1,702 citations, 1,100 URLs audited | No, observational by declaration | Structured data associated with citation |
| Vendor case studies | One site, before and after | No | Lifts of 19.72 percent and similar |
Does schema markup help with AI search?
There is no public evidence that schema markup causes AI citations. Google's own documentation says no special schema.org structured data is needed to appear in AI Overviews or AI Mode. The largest controlled test found no uplift across 1,885 treated pages. Independent testing found a JSON-LD-only fact read by zero of five AI systems.
That is the honest state of the evidence as of 2026-09-09, and every clause in it is sourced below. It is worth being precise about what it does and does not say, because the two most common readings are both wrong.
It does not say schema is useless. Rich results, entity resolution in a knowledge graph, merchant feeds and voice surfaces are real and measured, and the evidence base behind them is older and better controlled than anything in the AI-citation argument. Shipping schema is still the right call for most sites.
It also does not say schema definitely does nothing for AI citations. A null result on a specific population is not a proof of absence. The largest controlled test studied pages that were already heavily cited, and its authors say plainly that it cannot speak to pages outside that set.
What it says is narrower and more awkward: the specific claim being sold, that adding markup causes an AI assistant to cite you more, has no controlled evidence behind it, and the studies quoted in its favour mostly could not have detected it either way. That is a claim about instruments, not about schema.
From the field
None of this is an argument against shipping schema
Rich results, knowledge-graph entity resolution, voice surfaces and merchant feeds are real, measured, and worth the engineering. Google's own case studies on rich results report double-digit click-through improvements, and that evidence base is older, larger and better controlled than anything in the AI-citation debate. The argument here is narrower and it is about one specific claim, that adding schema causes AI assistants to cite you more. That claim is the one with no controlled evidence behind it, and it is the one being sold hardest.
Our reading of the public evidence base, 2026-09-09
What does Google actually say about structured data and AI Overviews?
Google's AI features documentation, last updated 2025-12-10 UTC, states that you do not need to create new machine readable files, AI text files, or markup to appear in these features, and that there is no special schema.org structured data you need to add. The same page asks that your structured data match the visible text.
We read both lines off the live document on 2026-09-09 rather than quoting somebody else's quotation of them, and the second line is the one that gets dropped. Google is not saying structured data is irrelevant. It is saying two separate things in the same breath, and they point in slightly different directions:
- No special markup is required. There is no AI-specific schema type, no
AIContentobject, no file you can add that puts you in the consideration set. - Your existing structured data should agree with what a reader sees. This is a maintenance instruction. A
Productblock claiming a price the page does not show is a correctness problem, and it becomes a bigger one when a system is reconciling two descriptions of the same page.
Neither of those is "add schema to get cited." Both of them are compatible with schema mattering somewhere upstream, which is exactly the ambiguity the rest of this post is about.
The complication is that Google has said something that sounds different. Practitioners routinely quote an April 2025 line to the effect that structured data gives an advantage in AI Overviews. We were not able to reconcile the two statements against a single primary source this pass. What we can say precisely is that "gives an advantage" and "is required" are different claims, and that the live documentation answers the second one and not the first.
Worth knowing
Google's own guidance has moved, and almost nobody puts the two statements side by side
The line practitioners quote most often is from April 2025 and says structured data gives an advantage in AI Overviews. The line on the live developer documentation, last updated 2025-12-10, says the opposite thing about requirement rather than advantage: there is no special schema.org structured data that you need to add. Both can be true at once, because "helps" and "required" are different claims, and that distinction is exactly what gets lost when a vendor quotes the first and skips the second. The same document also asks you to make your structured data match the visible text on the page, which is a maintenance instruction, not a visibility lever.
Google Search Central, AI features and your website, read 2026-09-09
There is one further wrinkle worth flagging because it circulates widely and we could not verify it. Multiple practitioner accounts describe a Google guide containing a mythbusting section that lists structured data as an AI-search requirement alongside llms.txt files and chunking content for LLMs, under the heading of hacks that do not help. We searched for that primary document on 2026-09-09 and did not locate it. The AI features page we did read contains no such section. We are recording that as unverified rather than repeating it as fact, because a claim about what a vendor said belongs to the vendor's own surface and nowhere else.
Why the "53 percent of AI-cited pages have schema" number proves nothing
Correlational findings on this topic are real and consistently positive. Pages cited by AI are roughly three times more likely to carry JSON-LD than pages that are not, measured across 6 million URLs. That number is doing all the persuasive work in the market, and it survives no contact with the confound sitting underneath it.
The confound is not subtle and the people who published the correlation named it themselves. Schema markup lives on better-maintained, more technically sophisticated sites. Those same sites publish stronger content, earn more links, keep their pages current, and rank better in ordinary search. Every one of those is independently associated with getting cited.
So the 53 percent figure is consistent with two completely different worlds:
- World A. Schema causes citations. Sites with schema get cited more because of the schema.
- World B. Competence causes both. Sites that ship schema also do nine other things right, and those nine things get them cited.
A correlation cannot distinguish World A from World B. That is not a criticism of the number, it is what the number is. The only way through is an intervention: take pages that did not have schema, add it, and compare them against pages that did not get it.
Ahrefs did exactly that, and published both halves three months apart. The correlational study came first and found the 3x gap. Then they ran a second study designed to isolate the effect, which is the part of the sequence that almost nobody quoting the first study mentions.
What the largest controlled test on schema and AI citations found
Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched each one against 3 control URLs from different domains at similar pre-period citation levels, and measured citations 30 days either side of the treatment date. Google AI Overviews moved minus 4.6 percent, AI Mode plus 2.4 percent, ChatGPT plus 2.2 percent.
This is the best-designed public evidence that exists on the question, and it is worth walking through the method rather than the headline, because the method is the part you can copy.
How the treatment date was defined. They pulled several million URLs cited in AI Overviews, retrieved the HTML history for each from their own crawler database, labelled whether each URL contained a <script type="application/ld+json"> block, and found the date on which schema presence transitioned from false to true. That transition is the treatment date. Note what this buys: the treatment is dated from an observation of the page's actual bytes, not from a client telling you when they shipped.
How the control group was built. For each treated URL they selected 3 control URLs from different domains with similar pre-period citation levels that never added JSON-LD. Different domains matters. Controls drawn from the same site share every site-level shock the treated pages do, which quietly removes the variance you were trying to measure against.
What the difference-in-differences test is for. Citations across the whole of AI search were moving during the window. AI Overviews were contracting, AI Mode was expanding fast. A simple before-and-after would have measured the platform trend and called it a schema effect. Comparing the change in the treated group against the change in the control group cancels anything that hit both.
Every public study on schema and AI citations, ranked by whether it had a control
| Source | Sample | Has a control group | Headline result | Source |
|---|---|---|---|---|
| Ahrefs, 2026-05-11 | 1,885 treated pages, 4,000 matched controls | Yes, 3 matched controls per treated URL | AIO -4.6 percent, AI Mode +2.4 percent, ChatGPT +2.2 percent | published |
| searchVIU, 2025-12-02 | 8 tests, 5 AI systems, 1 purpose-built page | Yes, a visible-HTML baseline test | JSON-LD-only price found by 0 of 5 systems | published |
| Otterly, 2026-03-23 | 319 prompts, 7 platforms, 1 site, 3 months | Informal, competitors moved in parallel | Deltas ran both directions, mechanism called indirect | published |
| Kumar and Palkhouski, 2025-09-13 | 1,702 citations, 1,100 URLs, 3 engines | No, observational by the authors' own statement | Structured data among strongest citation associations | published |
| Search/Atlas, December 2024, via Search Engine Land | Not disclosed in the secondary reporting | Not disclosed | No correlation between schema coverage and citation rates | unknown |
| Vendor case studies, various | 1 site each, before and after | No | Lifts reported at 19.72 percent, 25 percent and similar | published |
as of 2026-09-09
Method: Each row was opened at its own URL and the design column was read out of the study's own method section, never inferred from its title or from another page's summary of it.
The verdict on their own result is the sentence worth pinning to the wall.
So, overall, we can't tell whether the schema did a tiny bit of good or nothing at all.
The number that should end most before-and-after arguments
Read raw, before controls, Google AI Mode citations on the treated pages grew by 43 percent. After the control group's own growth was subtracted, the schema-attributable figure was plus 2.4 percent, and that residual was statistically indistinguishable from zero. Same pages, same window, same data, two answers off by a factor of eighteen.
This is the whole argument in one comparison, and it is worth sitting with, because 43 percent is exactly the kind of number that gets a case study written about it.
If you had shipped schema on those pages, watched AI Mode citations climb 43 percent over the following month, and published that, every sentence you wrote would have been true. The pages had schema. The citations went up 43 percent. You would have been reporting an observation, not inventing one. And your conclusion would still have been wrong, because AI Mode was expanding for everybody and your pages were carried along with it.
Nothing about your integrity is at issue in that scenario. The failure is instrumental. You measured a real change and attributed it to the only thing you happened to be watching. A control group is not a statistical nicety, it is the only thing standing between "citations went up" and "our change made citations go up."
What the AI Overviews decline does and does not mean
Treated pages lost 4.6 percent more AI Overview citations than matched controls, and that gap is statistically significant, with the authors putting the odds of seeing it by chance at roughly 1 in 2,500. It is also small in absolute terms, around 12 daily citations on pages mostly receiving hundreds, and both groups were already falling.
The temptation here is obvious and it should be resisted in both directions. This is not evidence that schema hurts AI Overview citations, and it is not evidence that schema does nothing. It is a real, small, unexplained gap, and the study says so.
Three things have to be true simultaneously for the honest reading:
- The gap is real. A 1-in-2,500 chance result over thousands of URLs is not noise you can wave away.
- The gap is small. Twelve citations a day against a baseline in the hundreds is under 5 percent of the page's traffic on that surface.
- The cause is unidentified. Adding schema often coincides with other changes: a template rewrite, a technical fix, a recrawl. The study pools all schema types together and cannot separate the markup from what shipped beside it.
We flag this because it is the part of the study most likely to be quoted out of context in either direction over the next year. A vendor will use it to argue the study is unreliable. A sceptic will use it to argue schema is harmful. Neither is what a small significant gap with no identified mechanism supports.
Can an AI assistant actually read the JSON-LD on your page?
Not during a direct fetch. In eight extraction tests run in October 2025 across ChatGPT, Claude, Gemini, Perplexity and Google AI Mode, a product price that existed only inside the page's JSON-LD was returned by none of the five systems. The same page's visible HTML, visible Microdata and visible RDFa were read by several.
This is the cleanest experiment in the entire literature and it deserves more attention than it gets, because unlike a citation study it has an unambiguous ground truth. The design was a single purpose-built page selling a fictional product line, with prices deliberately placed in eight different locations:
| Test | Where the price lived | Read by |
|---|---|---|
| 1 | Visible HTML | ChatGPT, Gemini, AI Mode after indexing |
| 2 | JavaScript-rendered DOM | Gemini live, AI Mode and Perplexity after indexing |
| 3 | JSON-LD only, nowhere visible | Nobody |
| 4 | JSON-LD injected via JavaScript | Nobody |
| 5 | Hidden Microdata | Nobody |
| 6 | Visible Microdata | ChatGPT, Gemini |
| 7 | Hidden RDFa plus conflicting JSON-LD | Nobody |
| 8 | Visible RDFa | ChatGPT, Gemini |
The pattern is not about the format. Read the table by column rather than by row and it resolves instantly: every fact these systems retrieved was visible on the page. Every fact that was not visible was missed, in every markup format tested. JSON-LD, hidden Microdata and hidden RDFa all failed identically, which is what you would expect if the parser is simply extracting rendered text and never touching the structured layer at all.
Two limits on this finding, both stated by the people who ran it. First, it tests Phase 4 of the pipeline, the direct fetch, and says nothing about the indexing phase, where structured data is demonstrably extracted. Second, five systems on one page in one month is a small sample even when the result is 0 for 8.
An earlier study found something compatible with a broader panel: 6 of 7 AI search platforms could not fetch or correctly interpret schema markup when asked for it directly, with only one retrieving the correct JSON-LD, and one engine reporting a schema type that did not exist on the page at all. Two independent tests, different months, different designs, same direction.
Where a schema effect would have to live
The four places between your page and an AI answer, and which ones read JSON-LD
- Your pageHTML plus a JSON-LD block in the head
- Search indexExtracts structured data, demonstrably
- Retrieval layerPicks a candidate set for the question
- Direct fetchParses only visible content, tested
- The modelTokenises whatever text it is handed
- The answerCites whatever survived the two paths
- Your pageSearch indexcrawled, schema extracted
- Your pageDirect fetchlive fetch, schema ignored
- Search indexRetrieval layercandidate set assembled
- Retrieval layerThe modelpassages handed over
- Direct fetchThe modelvisible text only
- The modelThe answercitation selected
The tokenisation argument, and exactly what it does and does not show
The mechanism argument underneath the empirical one is that a JSON-LD block does not survive tokenisation as structure. A model splits "@type": "Organization" into tokens for "type" and "Organization", which are indistinguishable from those same two words in a sentence. If that holds, machine-readable markup is not machine-readable to a language model.
The demonstration is straightforward and you can reproduce it in a tokeniser in about a minute. Paste a JSON-LD block, colour the token boundaries, and watch the syntax that carries the meaning get shredded. The quotes, the colon and the @ become their own tokens. What is left is a bag of words in a slightly unusual order.
A follow-up experiment made the same point from the other end. A fictional company's address was placed only inside invented, invalid schema on a page, never in the visible text, and both ChatGPT and Perplexity returned it when asked. The popular read was "schema works, the model read the markup." The author's own read was the opposite and it is the correct one: the markup was invalid, so nothing was parsing it as markup. The model read the text sitting inside the braces, exactly as it would have read the same words in a paragraph.
Be careful how far you push this. The tokenisation argument is about the model, and the model is one component in a pipeline. It says nothing about whether a search index extracts your Product block, which it demonstrably does, or whether the retrieval layer in front of the model uses index-derived structure to assemble candidates, which nobody outside those companies can observe. This is the distinction that makes the whole debate go in circles, and it deserves its own section.
Model reads schema, or retrieval layer uses schema? Two claims, one argument
Most schema arguments are two people describing different components. The claim that a language model parses your JSON-LD is testable and the tests say no. The claim that a retrieval pipeline in front of the model benefits from index-extracted structure is a different claim, largely untestable from outside, and not refuted by any evidence above.
The clearest public statement of this comes from the pro-schema side, and it concedes more than most schema advocacy does.
You're right that schema doesn't help the transformer parse. The mechanism isn't comprehension, it's the retrieval pipeline that runs in front of the model.
That is a serious position and we do not think it is wrong. It also changes what the evidence can settle. Split the claim in two and score them separately:
- Claim 1: the model parses your markup. Testable from outside. Tested twice. Both tests say no during direct fetch.
- Claim 2: the retrieval layer benefits from structure the index extracted. Not directly observable. Consistent with Google extracting structured data at index time. Also consistent with the null result, because the Ahrefs study measured pages already inside the consideration set, where retrieval had already selected them.
The frustrating consequence is that Claim 2 is compatible with every piece of evidence in this post, including the null. That does not make it true. It makes it unfalsified rather than supported, and those are different states that get spoken about identically.
If you want to argue for Claim 2, the test that would move it is not another correlation. It is an intervention on pages that are not currently being cited, measuring whether schema helps them enter the candidate set. The one large controlled study explicitly could not run that test, because every page in its sample already had 100 or more AI Overview citations before the treatment. That is the open question, stated precisely, and it is available to anyone with a corpus of uncited pages and the patience to do it properly.
Query fan-out, the explanation that competes with all of this
There is a mechanism that would produce most of the patterns in this post without schema doing anything, and almost nobody writing about structured data engages with it. Google's own documentation on AI features describes a query fan-out technique: the system takes one user question, generates a set of related sub-queries, runs searches across them, and assembles an answer from a wider set of supporting pages than a single query would have returned.
Read that against the correlational finding and something uncomfortable happens. If the systems that cite pages are running many sub-queries and pooling the results, then the pages that get cited are disproportionately pages that rank for many related things at once. That is a description of topical depth and site strength. It is also, separately, a description of the kind of site that ships structured data, because both are downstream of having an engineering team, a CMS with a template layer, and somebody whose job includes technical SEO.
Fan-out gives you a complete causal story in which schema is a passenger. Sites with resources rank for more sub-queries. Ranking for more sub-queries means appearing in more of the pooled candidate sets a fan-out assembles. Appearing in more candidate sets means more citations. Those same sites ship schema, because shipping schema is cheap and their stack makes it cheap. The correlation is real, reproducible at six million URLs, and causally inert.
We are not claiming fan-out is the true explanation. We are pointing out that it is a live alternative that the evidence base cannot currently rule out, and that a page recommending schema on the strength of a correlation is obliged to say so.
What fan-out predicts that a schema effect does not
The two stories are distinguishable, which is what makes this worth taking seriously rather than filing as a caveat.
A schema effect predicts that adding markup to a page should improve that page's citation rate, holding its ranking profile fixed. That is exactly the intervention the largest controlled test ran, and it netted 2.4 percent in AI Mode after controls, against 43 percent before them.
A fan-out effect predicts something different. It predicts that the lever is coverage, not markup: a page that starts ranking for six adjacent sub-queries instead of two should see its citation rate move, whether or not a single line of JSON-LD changed. It further predicts that the effect of schema should be roughly zero once you control for the ranking profile, which is what the controlled test observed.
Both the correlational finding and the null result are exactly what a fan-out story predicts. That is not proof. A theory that predicts everything you have already seen is cheap. But when the same evidence is equally consistent with two mechanisms, and the industry has been quoting it in support of only one of them, that is worth stating plainly.
Why this changes what you would test
If fan-out is doing the work, the intervention worth running is not a schema test at all. It is a coverage test: take a set of matched pages, expand one arm to answer a wider set of adjacent sub-questions on the same URL, freeze everything else including markup, and measure whether the treated arm enters more candidate sets.
That is a harder test to run than a schema test, which is probably why nobody has run it. Adding a JSON-LD block is a template change you can ship to 1,885 pages in an afternoon and verify with a curl. Expanding topical coverage is editorial work, it takes weeks, and it changes the page in ways that are difficult to hold constant. The easy test got run first because it was easy, and the industry read its subject matter as the important variable rather than as the convenient one.
We think that is the single largest structural bias in the current AEO evidence base. The variables that get tested are the ones that are cheap to manipulate, and the variables that are cheap to manipulate are rarely the ones with the largest effects. Schema is a template edit. Coverage, authority and entity clarity are not. Guess which one has a literature.
Reading vendor claims through this lens
The practical use is as a filter on other people's numbers. When a case study reports that citations rose after a schema rollout, ask what else moved on that site in the same window, and specifically ask whether the pages got more thorough as well as more marked up. The typical structured-data project is not a lone JSON-LD block landing on an otherwise frozen site. It arrives inside a technical SEO engagement, alongside internal linking work, template changes, page speed fixes and a content refresh.
Every one of those co-changes moves the ranking profile, which is the input a fan-out mechanism actually consumes. The schema is the most visible and most describable part of the release, so it gets the headline, and the parts that plausibly did the work are described as supporting activity. That is not dishonesty. It is what happens when the release note is written by whoever is proudest of their component.
What the academic paper everyone cites actually claims
The most-cited academic source on this question audited 1,100 URLs and 1,702 citations from Brave Summary, Google AI Overviews and Perplexity, using 70 product-intent prompts, and found structured data among the pillars most strongly associated with citation. The word doing the work in that sentence is "associated," and the paper's own abstract says so.
The study is observational and focuses on English language B2B SaaS pages; we discuss limitations, threats to validity, and reproducibility considerations.
We are not quoting that to dunk on the authors. The opposite. Declaring the design honestly in the abstract is exactly what a paper should do, and it puts the paper ahead of most of the vendor content citing it. The problem is entirely downstream: an observational finding gets quoted in an argument about causation, and the qualifier gets dropped somewhere between the abstract and the LinkedIn carousel.
Two further limits are worth carrying, both also stated by the authors:
- The corpus is English-language B2B SaaS pages. Whether the associations hold on documentation, on ecommerce, or on developer tooling is not established by this data.
- Structured Data is one pillar of sixteen in a composite quality score, and the paper's own strongest finding is that overall page quality predicts citation. That is the World B confound again, arriving from a completely different direction.
One circulating figure we deliberately do not quote: a correlation coefficient of 0.63 is widely attributed to this paper in practitioner threads. We read the abstract directly on 2026-09-09 and it is not there. It may well appear in the body of the paper. We did not verify it this pass, so we are not repeating it, and neither should the next person quoting us.
Why "we added schema and citations went up" is not evidence
Four things change the day your schema ships, and only one of them is your schema. The page gets edited, which makes it fresher. It gets recrawled, which re-enters it into ranking. The platform continues whatever trend it was already on. And whatever else was in that release ships alongside. A before-and-after credits all four to the one you were watching.
This is worth being concrete about, because "you need a control group" is advice everyone nods at and nobody follows. Here is the specific damage each confound does.
The platform trend. Already covered above and it is the biggest one. AI Mode grew for everyone during the study window. Any page measured before and after would have reported growth. Ahrefs' own difference-in-differences analysis turned a raw plus 43 percent into plus 2.4 percent by subtracting what the control group did anyway.
The freshness event. Adding schema is a page edit. Edited pages get recrawled, and freshness is independently associated with citation. If you add schema and citations rise, you have confounded markup with recency, and the only way to separate them is a control group that gets updated on the same day without the markup.
Co-occurring changes. Schema rarely ships alone. It arrives in a sprint with internal linking fixes, a template change, or a content refresh. The Ahrefs authors name this limitation on their own study explicitly: pages that add JSON-LD often change other things at the same time, and the pooled analysis cannot separate them.
Run-to-run volatility. The one almost nobody accounts for, because it is invisible unless you go looking. It gets its own section below.
Markup the crawler never sees. Also its own section. If your JSON-LD is injected client-side, your "treatment" was never applied at all, and you have run a test where the treatment arm and the control arm are identical.
The five confounds that eat a schema test, and the control that removes each one
| Confound | How it manufactures a lift | The control that removes it | Source |
|---|---|---|---|
| Platform trend | Every page moved, so your page moved | Matched control pages measured over the same window | derived |
| The freshness event | Adding schema is itself a page update | Control pages updated on the same day, schema withheld | derived |
| Co-occurring changes | Links, copy and technical fixes ship together | Freeze every other edit inside the test window | derived |
| Run-to-run volatility | Identical queries return different citations | Repeat each query enough times to see the noise band | measured |
| Invisible markup | Schema injected by JavaScript that crawlers never run | Fetch the raw HTML and grep for application/ld+json | published |
n = 12 · as of 2026-09-09
The uncomfortable implication is that a schema test with none of these controls is not a weak test. It is not a test. It produces a number, and the number is a measurement of something, but you have no way to say what. This is the difference our own measuring answer engine optimization lift work exists to close, and it is why we will not quote a lift figure to anyone.
Why do identical AI queries return different citations?
Because the systems are not deterministic and the retrieval step re-runs on every call. We measured this on our own query set before running anything else: 12 money queries, 5 identical repeats each, one model, one day, no intervention applied. 60 of 60 calls succeeded, and 7 of the 12 queries flipped outcome across identical repeats.
We call this the step-zero run because it comes before every other measurement, and because skipping it is what makes the rest of the numbers unreadable. The full method and the raw figures are published on why a single AI visibility score is noise, which is where the design lives rather than a summary of it.
Sit with what a 7-in-12 flip rate does to an uncontrolled schema test.
Suppose you track 12 queries, add schema, and re-run them. Five queries now cite you that did not before. That looks like a substantial win: a 42 percent improvement on your tracked set. Except we already know that running the same 12 queries twice, with nothing changed at all, produces movement on 7 of them. Your five wins are inside the range that nothing produces.
This is not a property of our particular query set being unusually unstable, as far as we can tell, though we would need a much larger sample to say that with confidence. It is a property of how these systems work. A retrieval step samples a candidate set. A generation step chooses what to cite from it. Neither is a lookup table, and neither promises to return the same answer twice.
There is a practical consequence that has nothing to do with schema. Any AI visibility dashboard reporting a week-over-week change on a small query set is reporting mostly noise, and it will keep doing so no matter how good its data collection is, because the instability is in the thing being measured rather than in the measurement. That is the argument in AI visibility tool versus AEO agency proof, and it is why we do not sell a score.
What does "no effect" actually look like? The null distribution
We ran 20,000 random splits of the same query set with no intervention applied at all, recording the difference between the two groups each time. That distribution is what "nothing happened" looks like on this query set. It centres on zero, mean plus or minus 0.0016, with a standard deviation of 0.215.
The reason to build this is simple and slightly humbling: without it, every observed movement reads as signal by default. You get a number, the number is not zero, and there is nothing to compare it against except your prior expectations. A null distribution replaces the prior with a measurement.
Read it as a ruler. If a genuine treatment produces a difference of 0.2 on this query set, and random splits of untreated data routinely produce differences of 0.215, the treatment has produced something smaller than the noise. That is not a small effect you detected. It is an effect you could not have detected, and the two get reported identically by anyone who has not built the ruler.
No page ranking for this query publishes one. We checked all ten on 2026-09-09. Not one carries a statement of what no effect looks like on the data they are showing you. That absence is not laziness; building a null distribution takes a query set, an outcome definition, and a few thousand shuffles, and most people writing about schema are writing an implementation guide rather than running an experiment. But it means every number on that page one is being read against an implicit null of exactly zero, which is the wrong ruler.
What is a permutation test, and why does it matter for AI visibility claims?
You take your own data, shuffle it into two random groups thousands of times with no treatment applied, and record the difference each time. The spread of those differences is your null distribution. Any real effect has to be larger than what shuffling produces by chance. It needs no assumptions about the shape of the underlying data.
The appeal for this specific problem is that it sidesteps a question nobody can answer. A conventional significance test assumes a distribution for your outcome. Nobody knows the right distribution for "did an answer engine cite this domain on this query on this run," and the honest answer is that it is a weird, lumpy, query-dependent thing. A permutation test does not care. It builds the reference distribution out of your actual observations.
The procedure, in the order you would run it:
- Define the outcome as a binary per observation. Cited or not cited, per query per repeat. Resist the urge to build a composite score at this stage; a score smuggles in weighting decisions you will not be able to defend later.
- Record your real difference. Treated group rate minus control group rate. One number.
- Discard the labels. Pool every observation, then randomly re-split into two groups of the same sizes as the real ones.
- Record the difference for that random split. Repeat several thousand times. We use 20,000, which is more than necessary and cheap.
- Count how often the shuffled difference is at least as extreme as the real one. That proportion is your p-value, and it is computed rather than looked up.
What you get out is a sentence you can defend under pressure: a difference this large appeared in N of 20,000 random splits of our own untreated data. That sentence survives a hostile reading in a way that "citations went up 19 percent" does not.
How small a citation lift can your test actually detect?
Our 12-query, 5-repeat design has a minimum detectable lift of 51.1 percentage points, roughly four times underpowered for the effects the market claims. We publish that number with the design, because a result below your own detection floor is not a small finding. It is no finding, wearing a number.
The minimum detectable effect, or MDE, is the smallest true difference your design has a reasonable chance of distinguishing from zero. It falls as your sample grows and rises with the variance in your outcome. On AI citation data, both terms work against you: the sample is expensive because every observation is an API call, and the variance is enormous because of the flip rate documented above.
Here is what that means concretely, and it is the sentence we would most like this post to be quoted for. A vendor case study reporting a 19.72 percent lift in AI Overview visibility, measured on one site with no control group, is not reporting a small effect that our instrument is too crude to see. It is reporting a number that its own design could not distinguish from nothing. The problem is not the size of the claim. It is that the design has no detection floor at all, because with no control group there is nothing to compute one against.
What your design size buys you, in minimum detectable lift
| Design | Total observations | Minimum detectable lift | What it can and cannot see | Source |
|---|---|---|---|---|
| 12 queries times 5 repeats | 60 | 51.1 percentage points | Cannot see any effect anyone has claimed | measured |
| 40 queries times 40 repeats | 1,600 | 9.9 percentage points | Can see a large effect, still blind to a small one | measured |
| A single before-and-after, no control | Undefined | Not computable | Cannot separate your change from anything else | unknown |
n = 60 · as of 2026-09-09
Method: Minimum detectable lift computed on the same query set and outcome definition as the step-zero run, published with the design rather than reconstructed for this post.
Scaling helps, and the arithmetic is worth knowing before you budget a test. Moving from 12 queries by 5 repeats to 40 queries by 40 repeats takes total observations from 60 to 1,600 and brings the floor from 51.1 percentage points down to 9.9. That is a 27-fold increase in sampling cost to buy a 5-fold improvement in resolution, and the endpoint is still a design that cannot see a 5-point effect.
We state this plainly because it cuts against our own interest. Our published instrument is not yet good enough to settle the schema question, and we are not going to pretend otherwise to sell a measurement engagement. The honest offer is a design that will tell you truthfully whether something moved beyond the noise, plus the number that says how small a movement it would have missed.
Reading the vendor case studies against a detection floor
Vendor case studies on schema are not fabricated. They are single-site before-and-afters with no control group, which is a design that cannot separate a treatment effect from a platform trend, a freshness event or ordinary volatility. The numbers are real observations. The causal claim attached to them is unsupported by the design that produced them.
The ones circulating on this SERP as of 2026-09-09 include a 19.72 percent increase in AI Overview visibility on a schema vendor's own site, a 25 percent increase in clicks and 30 percent increase in impressions for one client, and a hallucination-reduction claim for another. None discloses a control group, a sample size, or a measurement window.
Design quality
The same question, asked five ways, with five different amounts of evidence behind the answer
| A controlled test | A vendor case study | What we run | |
|---|---|---|---|
| Control group | Matched pages that got no schema | None | Held-out pages matched on the confounds |
| Repeats per query | Not applicable at page scale | One observation | Enough to establish the noise band first |
| Query set fixed in advance | Yes, the page set is the unit | Chosen after the result | Pre-registered before the treatment ships |
| Null distribution published | Reported as a margin of error | Never | 20,000 permutation splits, published |
| Detection floor stated | Implied by the confidence interval | Never | Stated in percentage points, with the caveat |
| What it can claim | No uplift detected at this scale | Something changed on one site | Nothing yet, and the post says so |
Apply the test from the section above to any of them. Ask three questions in this order:
- What did the pages that did not get schema do over the same window? If the answer is "we did not track any," the study has no denominator and the number is uninterpretable, whatever its size.
- What else shipped that month? Schema rollouts arrive with template changes and content refreshes. If the answer is "several things," you are reading a release note, not an experiment.
- How small a lift could this design have detected? If there is no control arm, the answer is that the question is not computable, and that is a stronger criticism than "the sample was small."
Note that the honest version of this exists and it is not hard to spot. Otterly's intervention study reports brand-coverage deltas running in both directions across seven platforms over the same window, and its own authors attribute the large AI Overview move to traditional-search mechanics because competitors who changed no schema moved in parallel. That last clause is an informal control group doing its job, and it is the single most methodologically honest sentence on the current page one. Almost nobody quoting that study mentions it.
The confound nobody checks: markup your crawler never sees
If your JSON-LD is injected by Google Tag Manager or any client-side JavaScript, the crawlers that matter here may never execute it. In a practitioner audit of 50 sites, 34 of them injected schema client-side, and none of those had schema detected by an AI crawler, against 92 percent of server-rendered ones. Their test arm and control arm were identical.
The check nobody runs first
What an AI crawler sees when your schema ships through a tag manager
Schema present, validator green
In the rendered DOM
Nothing
In the raw HTML
50
Sites audited
34 of 50, 68 percent
GTM-injected
0 percent
Detected, GTM-only
92 percent
Detected, server-side
- Inspect Element shows the schemaThis is the rendered DOM and it is the wrong instrument
- View source shows the schema
- curl the URL and grep for application/ld+jsonIf it is not in the raw bytes, the crawler that does not run JavaScript never sees it
This is a measurement problem before it is an implementation problem, and it is the reason it belongs in this post rather than in a setup guide. If you run a schema test on a site that injects markup through a tag manager, your treatment was never applied. You will get a null result, correctly, and draw the wrong conclusion from it.
The check takes one command and most people run the wrong instrument:
Inspect Element shows the rendered DOM, after JavaScript has already run, so a tag-manager injection looks present there and is still invisible to a bot that never executes the script. Read the bytes the server actually sent instead:
curl -sS -A 'Mozilla/5.0' https://example.com/your-page \
| grep -c 'application/ld\+json'
If that returns 0, your schema does not exist as far as any crawler that does not run JavaScript is concerned. Note the grep -c there counts matching LINES, not occurrences, so on a minified page it undercounts; use grep -o ... | wc -l if you need the true count rather than a presence check.
There is a wrinkle worth stating, because the picture is not uniform. Indexing crawlers and direct-fetch parsers behave differently. searchVIU's tests found that both Google AI Mode and Perplexity retrieved a JavaScript-rendered price after indexing, which means their indexing crawlers do execute JavaScript, while ChatGPT, Claude and Perplexity could not capture JavaScript content during a live fetch. So a client-side injection is not universally invisible. It is invisible to a specific and important subset, and it makes your test's treatment status a variable rather than a constant.
Does schema markup affect ChatGPT citations differently from Google AI Overviews?
The controlled test reports different per-platform numbers, plus 2.2 percent for ChatGPT, plus 2.4 percent for AI Mode, minus 4.6 percent for AI Overviews. Only the AI Overviews figure was statistically distinguishable from zero, and its authors decline to attribute it to schema. Read that as no reliable difference rather than as three different effects.
It is tempting to build a per-platform strategy on those three numbers, and the temptation should be resisted for a specific reason: two of the three are explicitly reported as indistinguishable from zero, which means the correct summary of them is "we did not detect anything," not "plus 2.2 and plus 2.4." Differencing two numbers that are each indistinguishable from zero produces a difference that is also indistinguishable from zero, and stacking uncertain estimates is how a per-platform playbook gets built out of nothing.
That said, there is a real architectural reason to expect platforms to differ, and it is the one from the pipeline section rather than anything in the citation data:
- Google AI Overviews and AI Mode retrieve from Google's index, which demonstrably extracts structured data. If a retrieval-side schema effect exists anywhere, this is where it would be.
- ChatGPT in its direct-fetch mode parses visible content and, per the extraction tests, misses JSON-LD entirely. A schema effect here would have to arrive through its search partner's index rather than through the page.
- Perplexity searches its own index first and, in the same tests, had the most selective extraction of any system, finding one of eight prices.
So the honest per-platform statement is architectural, not empirical: the places where a schema effect could exist are the indexed paths, and those are exactly the paths you cannot observe from outside. Nobody has measured a per-platform schema difference that survives a control group, and anyone selling one is extrapolating from the same three numbers you just read.
What the ten pages ranking for this query actually contain
We read all ten United States results for "schema markup for ai search" on 2026-09-09, page by page, and measured their body word counts rather than estimating them. Nine carry no original measurement. One does. Two put causality in the title and run no comparison in the body. The median comparable article is around 2,177 words.
This matters more than a competitive-audit section usually does, because on this particular question the SERP is the evidence base for most people. A practitioner asking whether schema helps reads three of these pages and forms a view. So it is worth stating what is in them.
| Rank | Page | Body words | Original measurement |
|---|---|---|---|
| 1 | SEOptimer | 2,177 | No |
| 2 | WordStream | 2,805 | No |
| 3 | Evertune | 1,033 | No |
| 4 | Wix Studio | 2,101 | No, vendor case studies |
| 5 | TheHoth | 2,178 | No |
| 6 | CMI Media Group | 753 | No, service page |
| 7 | YouTube video result | not comparable | Not applicable |
| 8 | Quora thread | not comparable | No |
| 9 | Schema App | 1,728 | No, vendor |
| 10 | LinkedIn post | not measured | No |
Word counts measured with a content scraper on main content only, markdown stripped of code fences, images and link URLs, alphabetic tokens counted. Nine of ten measured; one returned an HTTP 429 and is recorded as unmeasured rather than as zero, because those are different things.
The single largest gap is rank 3. Evertune's page is titled "Schema vs. No Schema: Does Structured Data Matter for AI Search?" and it does not run that comparison. It cites unnamed recent experiments with no design, no sample size and no numbers. It is the shortest credible page on the SERP at 1,033 words. A reader arriving on that title has asked precisely the right question and is handed framing.
To be fair to it, that page does concede something the pro-schema pages usually do not: that large language models may strip schema markup during tokenization. That single clause is the mechanism argument, sitting unremarked in the middle of a page whose title promises an experiment.
Rank 5 is the same shape, newer. TheHoth's page entered the top ten after the 2026-09-02 capture, and it is the closest title-level competitor to this post. We read it on 2026-09-09: 12 headings, and zero occurrences of methodology, sample size, experiment or baseline anywhere in the body.
Rank 4 and rank 9 are vendor-authored and both carry numbers that read as evidence and are not. The 19.72 percent AI Overview visibility increase quoted from Wix Studio's page was measured on the schema vendor's own site with no control group. Schema App's page is the oldest on the SERP and is sourced from conference keynotes and quoted chatbot responses.
The two most useful pages are the ones without a strong claim. WordStream's guide is the longest credible article on the page at 2,805 words and is a genuinely good type-by-type setup walkthrough with a validation workflow. It does not claim a citation effect. Search Engine Land's piece frames schema as infrastructure rather than a lever and states plainly that no peer-reviewed studies exist on this question. That was in print in March 2026, two months before the controlled test agreed with it.
How to read a schema study in five minutes
Five questions, asked in this order, sort a study from a story faster than reading the whole thing. What was the population. Was there a control group. What else changed. How many observations. And what would this design have failed to detect. A study that cannot answer the fifth has not stated a detection floor, which means its null result and its positive result are equally uninterpretable.
Run them against any page making a schema claim and the sorting is fast.
1. What was the population? Not the sample size, the population. Pages already heavily cited behave differently from pages nobody has ever cited, and a finding on one says almost nothing about the other. This is the limit the Ahrefs team put on their own study explicitly, and it is the single most consequential caveat in the entire evidence base, because the pages most companies care about are the uncited ones.
2. Was there a control group, and where did it come from? Ask specifically whether controls came from different domains. Controls drawn from the same site share every site-level shock and quietly cancel the variance you were trying to measure. The controlled test used 3 control URLs per treated URL, drawn from different domains, matched on pre-period citation level.
3. What else changed in the window? Schema rarely ships alone, and the studies that acknowledge this are more trustworthy than the ones that do not. A page that lists its own co-occurrence problem is telling you it looked.
4. How many observations, and over what window? Thirty days after treatment is a choice, and a slow-burn effect would be missed by it. So would a fast effect that decays. Neither is a flaw, but a study that does not tell you its window is not letting you weigh it.
5. What would this design have failed to detect? The detection floor. This is the question that almost never gets asked and it is the most powerful of the five, because it applies symmetrically. A null result from an underpowered design is not evidence of absence, and a positive result below the same floor is not evidence of presence. Both are the design failing to resolve, reported as a finding.
Apply that fifth question to our own published work and it fails: 51.1 percentage points is a floor that could not have detected any effect anyone has claimed, and we say so on why a single AI visibility score is noise rather than in a footnote. A checklist you will not apply to yourself is a rhetorical device.
Why every schema study you can find says schema works
Run the search yourself and you will notice something before you read a single methodology section. Ten pages rank for this query. Nine of them argue that schema helps. Not one publishes a negative result.
That is not what a healthy evidence base looks like on a genuinely open question. When a real effect is small and hard to measure, the honest literature is mixed: some studies find it, some do not, and the disagreement is the visible surface of the difficulty. A field where every published result points the same direction is usually a field where the results that pointed the other way did not get published.
There are three specific reasons a null result on schema is unlikely to reach you, and none of them requires anyone to lie.
Nobody publishes an experiment that found nothing
A structured-data vendor that ran a careful test and found no effect has produced a document that argues against its own product. It does not get written up. The same test with a positive result becomes a case study, a webinar and a conference talk. This is ordinary publication bias, the same mechanism that inflated effect sizes across nutrition and psychology for decades, and there is no reason to expect marketing research to be immune to it. If anything the incentive is sharper here, because the researcher and the seller are usually the same organisation.
Look at who authored the ranking pages. Two are by structured-data vendors. One is an agency service page in blog clothing. The single most-quoted figure on the SERP, a 19.72 percent AI Overview visibility lift, is a vendor reporting an uncontrolled before-and-after on its own website. None of that makes the number false. It does mean the number arrived through a filter that only passes results of one sign.
An uncontrolled test almost cannot return a null
This is the more interesting reason, and it does not require any bias at all. Recall the flip rate: 7 of 12 queries changed outcome across five identical repeats with nothing done to them. Now imagine a hundred teams each run an honest, uncontrolled before-and-after on a small query set after shipping schema.
Because the noise band is large and roughly symmetric, those hundred teams will get a spread of results. Some will see citations fall, some will see no change, and a substantial fraction will see a rise large enough to look like a win. The teams that saw a rise write a blog post. The teams that saw a fall assume they implemented it wrong. Nobody in either group has run a test capable of distinguishing signal from the flip rate, so the sign of their result was decided by noise, and only one sign gets published.
The output of that process is a literature of positive case studies with no shared methodology, which is precisely what page one of this query looks like. The volatility that makes individual results meaningless is the same volatility that guarantees a steady supply of impressive ones.
The measurement window rewards optimism
The third reason is timing. A team ships schema and measures over the following weeks, which is the same window in which the pages were recrawled, in which whatever else shipped in that release took effect, and in which the platform itself kept moving. AI Overview and AI Mode surfaces have been expanding across most of the period these case studies cover. A rising platform baseline means an uncontrolled after-measurement starts with a tailwind that has nothing to do with the intervention.
A held-out control arm removes this entirely, which is why the one large controlled study saw 43 percent collapse to 2.4 percent when controls were applied. Most published tests have no control arm, so most published tests are measuring the platform and attributing it to themselves.
What to do with this
Treat the uniformity of the literature as information about the literature rather than about schema. When every available study agrees and none of them shares a design, the agreement is a property of the filter, not of the world.
The practical version is a single question to ask of any schema result someone shows you, including ours: would this study have been published if it had come out the other way? If the answer is no, the result carries much less weight than its confidence interval suggests, and it very likely does not have a confidence interval at all.
We publish our null partly because it is what we measured and partly because a market with no negative results is a market where nobody can calibrate. The 51.1 point detection floor on our own design is not a flattering number to put in writing. It is in writing because a post arguing that other people should state their limits, while omitting its own, would not be worth reading.
Objections to this post, answered
The strongest objections to the argument here come from people who have looked at the same evidence, and three of them are good. We would rather state them properly than build weak versions to knock down, because a methodology post that cannot survive its own critique has the same problem as the case studies it objects to.
Objection 1: "You are quoting a null result on pages that were already cited. That is not the interesting population."
Correct, and this is the strongest objection by a distance. Every page in the controlled study already had 100 or more AI Overview citations before schema was added. Those pages were already inside the consideration set, already being crawled and surfaced. A finding that schema does not push an already-visible page higher is genuinely silent on whether schema helps an invisible page become visible.
Our answer is not to dismiss it. It is that the claim being sold does not carry that qualifier either. Nobody's pitch says "schema will help your uncited pages enter the candidate set, and will do nothing for pages already being cited." The pitch is undifferentiated, and the one controlled test we have refutes the half of it that was testable. The other half is open and we said so twice above.
Objection 2: "Tokenisation is about the model, and the model is not where retrieval happens."
Also correct, and we gave this its own section rather than burying it. If a schema effect exists, the indexed path is where it would live, and Google demonstrably extracts structured data at index time. The extraction tests measure direct fetch and say so.
Where we push back is on what that buys the argument. An unfalsifiable location for an effect is not evidence for the effect. "It happens in a layer you cannot observe" is a hypothesis, and it has the same epistemic standing as any other untested hypothesis, which is to say it is worth testing and not worth selling. The design that would test it is in the protocol section, and it is available to anyone.
Objection 3: "Three independent studies find a positive association. You are cherry-picking the one null."
The association findings are real and we cite them rather than dodging them: an observational audit of 1,702 citations across 1,100 URLs ranking structured data among the strongest citation associations, a correlational finding that AI-cited pages are around three times more likely to carry JSON-LD, and a meta-analysis synthesising 54 studies that reportedly ranks structured data as a positive factor while noting limited evidence that engines can see schema when searching.
The answer is the one from the correlation section, and it is not a rebuttal of those studies so much as a statement about what they can bear. Every one of them is observational. Not one has a treatment and a control. They can establish that schema and citations co-occur, which is not disputed by anybody including the null result. They cannot establish direction. When the same organisation ran the correlation and then ran the intervention, the correlation was 3x and the intervention was zero, and that is exactly the pattern you would expect if competence, rather than markup, were driving both.
There is also a smaller point worth making about the meta-analysis specifically, since it is quoted often. Synthesising 54 studies that are individually observational does not produce causal evidence. It produces a well-summarised body of observational evidence, which is useful and is not the same thing.
Objection 4: "Your own instrument is too weak to have an opinion."
True, and we say so in its own section further down. Our design has a 51.1-point floor and we hold no lift result of any kind. What we hold is the noise measurement, and the noise measurement is what makes other people's numbers readable. You do not need a powerful instrument to point out that somebody else's instrument has no control group. Those are different claims requiring different evidence, and the second one is settled by reading the method section.
The freshness confound, in detail
Adding schema is a page edit, and edited pages get recrawled. If citations rise after your schema ships, you have confounded markup with recency, and separating them requires a control arm that gets edited on the same day without receiving the markup. Almost no published schema test does this, including the good ones.
This one deserves detail because it is the confound most likely to survive an otherwise careful design, and because it explains a specific pattern in the data.
The mechanism is straightforward. A recrawl re-evaluates the page. Freshness is independently associated with citation across every study that looks for it. So the sequence looks like this:
- You edit the page to add JSON-LD. The file's modification date changes.
- The crawler visits, sees a changed page, and re-processes it.
- The page re-enters retrieval consideration with a recent timestamp.
- Citations move.
Steps 2 through 4 would have happened if you had fixed a typo. The markup is along for the ride.
The controlled test handles this partly by design and partly by luck. Because its control pages were matched on pre-period citation level rather than on edit recency, the control arm was not systematically edited on the treatment date. The authors ran a fourth sensitivity check with a symmetrical window that excluded the recrawling period specifically to see whether the result was sensitive to how before and after were defined, and it was not. That is the right check and it is more than most tests do.
For a test you run yourself, the control is cheaper than it sounds. Edit the control pages on the same day. Change a word, update a date, touch anything that triggers a recrawl, and withhold only the markup. Now both arms received the freshness event and the only difference between them is the variable you care about. This is the third step of the protocol below, and it is the step people most often skip because it feels like doing pointless work.
The reason it is not pointless is that skipping it makes a positive result uninterpretable, which is the outcome you would most regret. A null result with a sloppy control is annoying. A positive result with a sloppy control is worse, because you will act on it.
The Citon Variable Isolation Protocol
The Citon Variable Isolation Protocol is the seven-step design we use to make a single on-page change measurable against AI citations. It exists because the alternative, a before-and-after on a live site, cannot separate a treatment from a platform trend, a freshness event or ordinary run-to-run volatility. Every step removes one specific confound.
We name it because the parts get quoted separately and drift apart. People take the control group and skip the repeats, or take the repeats and skip the pre-registration, and end up with a design that looks rigorous and closes none of the holes. The protocol is the whole set or it is not the protocol.
Step 1. Pre-register the query set. Write down the exact queries, verbatim, before you look at any data, and commit the file. Not a theme, not a topic, the literal strings. This is the cheapest step and the one that prevents the most self-deception, because a query set chosen after you see the results is a query set selected for the result.
Step 2. Establish the noise band before you treat anything. Run every query in the set N times with no change applied, and record how often each one flips. This is the step-zero run. If more than half your queries disagree with themselves, as more than half of ours did, you now know the size of the effect you would need before any observed change means anything.
Step 3. Split pages into treated and held-out arms, matched on the confounds. Match on pre-period citation rate first, because that is the variable most predictive of the outcome. Then on page type, publication date and update recency. Assign randomly within matched pairs so your own judgment about which pages "should" respond never touches the assignment.
Step 4. Apply the treatment, and only the treatment. Ship schema on the treated arm. Change nothing else on either arm for the duration of the window. If your release process cannot hold that line, the test cannot run, and knowing that in advance is worth more than a result you will not be able to defend.
Step 5. Verify the treatment actually landed. Fetch the raw HTML of every treated page and confirm the markup is in the bytes. This is the step that catches the client-side injection problem from the section above, and it takes one loop. A test whose treatment was never applied returns a null result for the wrong reason.
Step 6. Re-measure both arms with the same repeat count. Same queries, same N, same model, same day if you can manage it. Model version drift across a long window is its own confound and it is the one you have least control over, so compress the window rather than extending it.
Step 7. Publish the null distribution alongside the effect. Permute your own data a few thousand times, report how often the shuffle produced a difference at least as large as the real one, and state the minimum detectable effect of the design. An effect reported without a null is an assertion.
What the readout looks like
A two-arm schema test, read the way it has to be read
40 pages, schema shipped
Treated pages
40 pages, no change
Held-out control
plus 6.2 points
Treated change
plus 5.8 points
Control change
plus 0.4 points
Difference
9.9 points
Detection floor
- Query set pre-registered before the treatment shipped
- Every query repeated enough times to see the noise band
- No other page edits inside the window
- Result clears the stated detection floor0.4 points against a 9.9-point floor is not a small effect, it is no result
The result that panel shows is the ordinary one and it is the reason most teams do not run this design. A difference of 0.4 points against a floor of 9.9 is a real answer, it just is not a marketable one. The protocol's job is to make "nothing happened" a reportable outcome, because a method that can only produce wins is not measuring anything.
A worked example: running the protocol on a documentation site
Here is the protocol applied end to end on a plausible case, with the arithmetic done rather than gestured at. A developer-tools company with 180 documentation pages wants to know whether adding TechArticle and BreadcrumbList markup moves AI citations. The whole design fits on one page and the decision it produces is frequently "do not run this."
The setup. 180 doc pages. Roughly 40 of them get cited occasionally by at least one assistant. The team can afford about 2,000 assistant calls for the whole exercise. They want to detect a lift they would act on, which they define, after being pushed on it, as 10 percentage points.
Step 1, pre-registration. They write 40 queries verbatim into a file and commit it. The queries are the real buyer questions, phrased the way a buyer phrases them, which for this audience means things like "how do I paginate results from the X API" rather than "best API for X." That distinction matters more than it sounds: the average question a buyer puts to an assistant is far longer than a search query, and a query set written in search-keyword shape is measuring a different population than the one that buys. The question sets we hold for this audience are in the questions developer tool buyers ask AI.
Step 2, the noise band. They run all 40 queries 5 times each with nothing changed. That is 200 calls out of the 2,000 budget, spent before any treatment exists. Suppose 21 of the 40 queries flip at least once, which is roughly the rate we measured on our own set. They now know that more than half of their tracked queries move on their own inside a single day.
That number alone changes the conversation. A stakeholder expecting a clean before-and-after now has a concrete reason why the after will disagree with the before regardless of what ships.
Step 3, matched arms. They take the 40 pages with any citation history and split them into two arms of 20, matched on pre-period citation rate first, then on page type and last-updated date. Assignment is randomised inside each matched pair, by script, not by judgment. The moment somebody says "let us put the important pages in the treated arm" the test is over.
Step 4, the treatment. TechArticle and BreadcrumbList go on the 20 treated pages. Nothing else changes on either arm for four weeks. This is the step that kills most real tests, because a docs site with an active team cannot usually freeze for four weeks. If that is true, say so now and do not run it.
Step 5, verification. They fetch the raw HTML of all 40 pages and confirm the markup is in the bytes on exactly the 20 treated ones and absent on the 20 controls. Getting a non-zero count on a control page means the assignment leaked, and getting a zero on a treated page means the deploy did not land. Both are common and both are silent.
Step 6, re-measurement. Same 40 queries, 5 repeats, same model, same week. Another 200 calls. Running total: 400 of 2,000.
Step 7, the readout. Treated arm citation rate moved from 0.31 to 0.37. Control arm moved from 0.30 to 0.35. Difference in differences: plus 0.01, or one percentage point.
Now the part that decides everything. Their design is 40 queries by 5 repeats across two arms, which is 400 observations total, not the 1,600 that reaches a 9.9-point floor. Their actual detection floor lands well above their 10-point acting threshold. The correct conclusion is not "schema produced a one-point lift." It is "this design could not have detected the effect we care about, and it did not detect one."
Those are different sentences and only the second is defensible. The first would be quoted in a deck within a week.
What they should have done with the 2,000 calls. Run the noise band on 40 queries at 5 repeats, see the flip rate, and conclude before treating anything that 400 observations cannot resolve a 10-point effect on data this noisy. Then spend the remaining 1,600 calls on 40 queries at 20 repeats per arm, which gets substantially closer, or spend them on a question with a larger prior effect, such as whether the documentation is retrievable at all by a fetcher that does not run JavaScript.
That last option is the one we would actually recommend for most teams in this position, and it is worth saying plainly because it argues against selling a measurement engagement. If your docs are behind a client-side router, or your schema ships through a tag manager, you have a retrievability problem whose expected effect size is enormous and whose test costs one curl. Fix that first. The schema-causation question is interesting and the retrievability question is worth money.
One variation worth knowing. If the team's real question is about pages that are not currently cited, the design changes in one important way and gets harder. The matched pairs are matched on being uncited, the outcome becomes a binary entry event rather than a rate change, and entry events are rarer than rate changes, so the sample requirement goes up rather than down. That is the test nobody has run, ourselves included, and it is where the interesting answer lives.
Sizing the test before you spend anything on it
The two numbers that decide whether your test is worth running are the flip rate on your own queries and the size of effect you would act on. Establish both before you buy a single API call, because the arithmetic frequently says the test you had in mind cannot answer the question you had in mind.
Work it in this order:
- Run 10 to 15 queries 5 times each with no treatment. This is cheap, it is one afternoon, and it gives you your flip rate. Ours was 7 in 12.
- Decide the smallest effect that would change what you do. If a 3-point lift would not change your roadmap, do not design a test that resolves 3 points. Most teams, pressed on this, name something between 10 and 20 points.
- Compare that threshold against what your budget buys. 60 observations bought us a 51.1-point floor. 1,600 observations gets to 9.9. If your acting-threshold is below the floor your budget reaches, the test cannot pay for itself, and the correct decision is not to run it.
Point 3 is the one worth internalising. A test that cannot detect the effect you care about is worse than no test, because it will return a number, and that number will be read as evidence by whoever sees it next. The 12-by-5 design we published is useful for exactly one thing, which is establishing that the noise band is large. It cannot be used to argue that schema did or did not work, and we say so in the source post.
There is a cheaper alternative worth naming honestly, since we would rather you skip a bad test than run one: on a small site, do not test schema at all. The evidence base says the expected effect is at or near zero, the cost of shipping schema correctly is low, and the value of a correctly-powered test would exceed the value of the finding. Ship the markup for the rich-result and entity reasons that are actually established, and put the measurement budget on something with a larger prior.
Running the checks that are worth it without a research budget
The protocol above assumes a team that can afford a control arm, a few thousand API calls and someone who enjoys permutation tests. Most sites reading this have twenty pages, no research budget, and a reasonable objection: none of this is available to us, so what are we supposed to do.
The answer is not a smaller version of the same experiment. A smaller version has a worse detection floor, and a test that cannot see the effect you care about returns a number anyway, which is the failure mode this whole post is about. The answer is to stop trying to measure the effect and start verifying the things that are cheap, binary and currently broken on a surprising number of sites.
Here are four checks, in the order of how often we find them failing. None requires a control group. All four are answerable in an afternoon.
1. Is the markup actually in the served HTML
Fetch your own page the way a crawler does, with no JavaScript execution, and look for the JSON-LD block:
curl -s https://example.com/your-page | grep -c 'application/ld+json'
A zero here means the treatment was never applied, and everything downstream of it is measuring nothing. This is the confound covered at length above, and it is the one we find most often, because client-side injected structured data looks perfect in a browser and perfect in most validators, which render the page first. A validator that executes JavaScript cannot tell you whether a crawler that does not execute JavaScript will see your markup. Only a raw fetch can.
2. Does the markup describe what a reader can actually see
Google's guidance asks that structured data match the visible content of the page. This is not a technicality. A page whose markup claims an aggregate rating, an author, or a FAQ that does not appear in the visible text is carrying a liability rather than an asset, and the enforcement risk is real in a way the citation upside is not.
Read your own JSON-LD next to your own page and check every claim resolves to something on screen. Where it does not, the fix is usually to delete the property rather than to add the content, because most of these fields were added by a plugin that was guessing.
3. Does it validate, and does it earn a result you want
Run the page through Google's Rich Results Test and the schema.org validator. These answer different questions: the first tells you whether the markup is eligible for a specific search feature, the second whether it is well formed. Eligibility for a rich result is the one benefit in this entire area that is documented, established and observable, so it is the one worth optimising for.
If a type earns you no rich result and describes nothing visible, it is doing nothing for you in either search or AI answers, and it still has to be maintained. Delete it.
4. Can you keep it correct for two years
This is the check nobody runs and the one that decides the outcome. Structured data breaks quietly. A template change drops a field, a CMS migration mangles the escaping, a product model changes and the markup keeps asserting the old shape. Nothing alerts you, because invalid structured data is not a page error, it is a silently ignored block.
Ask who owns the markup, whether it is generated from the same data the page renders from, and whether anything in CI would fail if it broke. Markup generated from the page's own data cannot drift from the page. Markup maintained by hand in a separate template will drift, and the only question is when. If the honest answer is that nobody owns it, ship fewer types and generate them properly rather than shipping many and maintaining none.
What this buys you
Not a citation lift. On the evidence in this post, you should expect approximately nothing on that front, and any of the four checks above is a better use of an afternoon than an underpowered experiment.
What it buys is a page whose machine-readable description of itself is accurate, served, valid, and maintainable. That is worth having on its own terms, it is the thing Google's documentation actually asks for, and it is the only part of this subject where the expected value is clearly positive. The measurement budget you did not spend on a 51 point detection floor is better spent on almost anything else.
Which schema types are actually worth shipping, and why
The useful list is the traditional search list: Organization, Article, Product, BreadcrumbList, LocalBusiness, and Dataset for research pages. No type has been shown to move AI citations specifically. Choose types that describe what is visibly on the page, because Google's own guidance asks that your structured data match the visible text, and a claim with no visible counterpart is a liability.
That answer is less exciting than the ranked type lists on the pages currently ranking for this query, and the reason is that those lists are not derived from anything. SEOptimer recommends nine types and WordStream walks through a type-by-type setup, and both are competent implementation guides. Neither measured which types help, because no public study has. The Ahrefs test pooled Article, FAQ, Product, HowTo and Organization together and says explicitly that separating them is future work.
So the honest basis for choosing types is not AI citations. It is these three, in order:
- Does it earn a rich result you actually want? Product, Review, Recipe, Event and Video still drive real, measured click-through improvements in ordinary search. That evidence base predates this whole argument and is stronger than anything in it.
- Does it describe an entity you want resolved correctly? Organization with
sameAspointing at profiles that already exist elsewhere is doing entity-disambiguation work, and the cost of getting it wrong is being merged with a similarly-named company. - Can you maintain it? Exotic types carry ongoing maintenance cost and Google periodically deprecates support. The relevant question is never "does this type help," it is "does this type help enough to keep correct for two years."
We would add one more, specific to research and documentation properties, in the next section.
Is FAQPage schema still worth adding for AI answers?
For rich results, mostly not. Google narrowed FAQ rich results in August 2023 to a small set of well-known authoritative government and health sites, and removed HowTo rich results on desktop entirely. The markup is still valid. The visible payoff for nearly every site went away.
For machine extractability, it can still earn its place, and the argument is narrower than the one usually made for it. A question-and-answer block is trivially easy for any parser to segment, and the structure of the content itself, short self-contained answers to explicit questions, is the shape that gets lifted into an answer regardless of whether the markup is read. The value is in writing the FAQ, not in wrapping it.
This distinction matters for testing, because FAQ schema is the type most people reach for when they want to test schema, and it is the weakest possible choice:
- Its rich result is already gone, so any traditional-search pathway is closed.
- It is the type most likely to describe content that is on the page in visible form anyway, so a parser reading rendered text already has it.
- A null result on FAQPage gets generalised to all schema, which is a measurement error rather than a finding.
If you are going to test one type, the current eligibility rules make FAQPage the one least likely to show you anything. Test a type whose rich result is live and whose content is not otherwise visible, or accept that you are testing the weakest case and report it as such.
We ship FAQPage on this post, and the reason is the extractability one rather than the rich-result one. Saying that out loud on a post arguing about schema seemed better than shipping it silently and inviting the obvious reply.
Dataset schema, the type no page ranking for this query uses
Every page on this SERP recommends schema types and none of them ships the one that fits a page carrying original measurement. Dataset describes a collection of data with a distribution, a variable measured, and a licence, and a post publishing a permutation test with 20,000 splits is a dataset by any reading of that vocabulary.
We use it here for our own runs, and the reasoning is worth stating because it is the only place in this post where we recommend a type on grounds other than "it is established":
- It matches the visible content. The numbers are on the page, in tables, with their sample sizes and dates. That satisfies Google's own match-the-visible-text instruction without any strain.
- It is a correctness claim we can actually stand behind. A
Datasetnode assertingvariableMeasuredand adistributionis a factual description of what the page contains, not a positioning statement. - We do not claim it earns citations. There is no evidence that it does. It is the honest markup for the content, which is a different and lower bar than a visibility lever, and it is the only bar this post is willing to defend.
That third point is deliberately deflationary. Recommending a type because it accurately describes your page is not a growth tactic, and dressing it up as one would be the exact move this post spends eleven sections objecting to. The relationship between how AI citations get picked and how your page is marked up is, on the current evidence, weaker than the market believes, and a type recommendation should not pretend otherwise.
What to measure instead of a single AI visibility score
Instrument five things and a score becomes unnecessary. The repeat-stability rate on your own queries. The null distribution from permuting your own data. The minimum detectable effect of the design you can afford. A held-out control set. And the query set itself, fixed in advance. A score compresses all five into one number and discards the only information that made it readable.
The case against a single visibility score is not that scores are bad. It is that a score has no error bar, and on data with a 7-in-12 flip rate the error bar is most of the number. We have written the long version of this on AI citation measurement, and what counts, and the short version is that a dashboard reporting a weekly delta on 20 queries is reporting a coin flip with a chart around it.
What the five measures buy you, individually:
- Repeat stability tells you whether any observed change is even in the range of interpretable. It is the cheapest measurement in this list and almost nobody runs it.
- The null distribution replaces "not zero, therefore something" with a calibrated comparison.
- The detection floor tells you in advance whether the test can answer the question, which is the only moment that information is actionable.
- A held-out control is the difference between a measurement and an anecdote.
- A fixed query set stops you selecting the queries that moved and calling them the result.
None of that requires a vendor and none of it requires us. It requires an outcome definition, a scripted query runner, and the discipline to write the query set down first. The reason it is rare is not cost. It is that the output is frequently "we cannot tell," and that is a hard thing to put in a board deck. The argument for putting it there anyway is in AI citations, prediction versus proof.
How this changes if you sell an API or a developer tool
Developer buyers ask AI assistants different questions, and the questions behave differently. Comparison and integration queries return citations skewed toward documentation, changelogs and community threads rather than marketing pages, which changes both what you would treat in a schema test and which pages you would hold out as controls.
The population problem from earlier is sharper here. The one large controlled test drew its sample from pages already carrying 100 or more AI Overview citations, and most developer-tool documentation is nowhere near that threshold. That means the null result is least applicable exactly where a developer-tool company would most want an answer, on pages that are not currently in the consideration set at all.
This is the open question we flagged above, and it is genuinely open. If schema helps anywhere, helping an uncited page enter a candidate set is the most plausible place, and nobody has run that test. We have not run it either. What we can say is what a test would need to look like, which is the protocol above with one change: the matched pairs are matched on being uncited rather than on pre-period citation rate, and the outcome becomes an entry event rather than a rate change.
For teams in this position, the sequencing we recommend is unglamorous:
- Fix retrievability before markup. If your docs are behind a client-side router that a non-JavaScript fetcher cannot read, no markup question matters yet.
- Write the answers in visible text. Every extraction test in this post says the same thing: visible content is what gets read.
- Ship correct schema, cheaply, and stop thinking about it. Then measure something with a larger expected effect.
The question sets that developer-tool buyers actually put to assistants are collected in the questions developer tool buyers ask AI, and the practice-level view is on developer tools and APIs.
How the evidence base actually moved, in order
The schema and AI citations question went from anecdote to a controlled null in about twenty months, and the sequence matters because most of the arguments still circulating are anchored at a point in it that has since been superseded. Six dated events and one gap.
TIMELINE
How the evidence base actually moved, dated
September 2025, the tokenisation argument
Mark Williams-Cook demonstrates that a JSON-LD block does not survive tokenisation as structure, splitting into tokens indistinguishable from ordinary words. Search Engine Roundtable reports it. The mechanism argument starts here.
September 2025, the observational paper
Kumar and Palkhouski publish an audit of 1,702 citations across 1,100 URLs, finding structured data among the strongest citation associations, and state in their own abstract that the study is observational.
October to December 2025, the direct-fetch tests
searchVIU runs 8 extraction tests across 5 AI systems on a purpose-built page. A price present only inside JSON-LD is returned by zero of them. Visible Microdata and visible RDFa are read; hidden markup of every format is not.
December 2025, Google's documentation settles the requirement question
The AI-features page states that no special schema.org structured data is needed to appear in AI Overviews or AI Mode, and asks that structured data match the visible text on the page.
March 2026, the first intervention study
Otterly ships schema on one site and tracks 319 prompts across 7 platforms for three months. Its deltas run in both directions and its own authors attribute the movement to traditional search, because competitors who changed nothing moved in parallel.
May 2026, the controlled test at scale
Ahrefs tracks 1,885 pages that added JSON-LD against 4,000 matched controls using difference-in-differences, and reports no uplift on any platform. The raw AI Mode figure of plus 43 percent collapses to plus 2.4 percent once controls are applied.
What is still missing, 2026-09-09
A test that isolates schema on pages not already inside the citation consideration set, with a pre-registered query set, repeats sufficient to establish the noise band, and a published null distribution. Nobody has run it, including us.
Two things stand out reading it in order. The first is that the mechanism argument arrived before the empirical one, which is unusual and helpful: the tokenisation demonstration in September 2025 predicted what the extraction tests found in October and the controlled test found in May. When a mechanism story and two independent empirical results point the same way, the burden shifts.
The second is that the field's most careful writers arrived here early and were largely ignored. The "infrastructure, not a magic bullet" framing, the December 2024 study finding no correlation between schema coverage and citation rates, and the plain statement that no peer-reviewed studies exist on this question were all in print in March 2026, two months before the controlled test confirmed them. The market kept selling the other thing.
Worth noting what has not been superseded: the observational paper still stands as evidence of association, the tokenisation post still stands as a mechanism argument about the model specifically, and the retrieval-layer position remains unfalsified. None of those has been knocked down. They just answer narrower questions than they are usually quoted for.
What would change our mind
A post that argues other people's evidence is unfalsifiable owes the reader its own disconfirmation conditions. Here are ours, stated specifically enough that somebody could go and satisfy them.
A pre-registered intervention on uncited pages showing entry. This is the big one and it is the test nobody has run. Take a corpus of pages that currently receive zero AI citations, split them into matched arms on their existing ranking profile, ship schema to one arm only, freeze every other edit including content and internal links, and measure entry into the citation set over a fixed window. If the treated arm enters at a meaningfully higher rate than the control arm, with the query set committed in advance and a stated detection floor, we would revise the central claim of this post. The largest existing study could not run this, because every page in its sample already had 100 or more citations before treatment.
A replicated direct-fetch result showing markup being read. Our extraction test placed a fact only inside JSON-LD and found 0 of 5 systems retrieved it. That is a small test on a single day. If somebody runs the same design across more systems, more facts and repeated days and finds that some engines do read markup on a direct fetch, the mechanism section above needs rewriting, and the practical advice changes for the specific engines involved.
A vendor case study with a control arm. We would take a single uncontrolled case study far more seriously if it carried a held-out set of matched pages from the same site that got no markup. No published case study we found does this. One that did, showing a treated-versus-control gap larger than the site's own repeat-measure noise, would be the strongest piece of applied evidence in the field, and would move us more than another six-million-URL correlation.
A stated flip rate near zero on somebody else's query set. Our volatility finding rests on 12 queries, 5 repeats, one model, one day. That is thin and we have said so. If several teams publish repeat-measure baselines on larger sets and find that identical queries return identical citations most of the time, then the noise-band argument that carries much of this post weakens considerably, and uncontrolled before-and-after tests become more readable than we currently think they are.
A replication of the null in a different vertical. The one large controlled test ran on a single corpus with a single content profile. Effects that are absent in one vertical are not guaranteed absent everywhere, and there are reasons to think markup could matter more where the entities are unusually well defined, such as local business, recipes, events or product catalogues, all of which have mature schema types and long-standing rich results attached. If a controlled intervention in one of those verticals found an uplift that survived matched controls, the honest conclusion would not be that this post was wrong, it would be that the answer is vertical-dependent and the general claim was overstated in both directions. We would rather be corrected that way than keep quoting one study as though it settled every case.
A first-party statement from a platform. If Google, OpenAI or Perplexity documented that structured data is consumed as a ranking or citation input for AI answers, rather than as an eligibility signal for search features, that would settle the mechanism question directly. Google's current documentation says the opposite, that no special structured data is needed for AI features, and it was updated on 2025-12-10.
What would not change our mind
Naming these is equally useful, because these are the arguments we expect to receive.
Another correlational study, at any sample size. Six million URLs did not settle it and sixty million will not, because the confound is site competence and no amount of observational data separates it from the treatment.
A larger uncontrolled before-and-after. Volume does not fix a missing control arm. A hundred sites each measuring themselves against their own past produces a hundred measurements of the platform trend.
An appeal to how the technology plausibly works. The tokenisation argument is a real mechanism argument about the model and we cite it as such. It is not evidence that citations move, and the retrieval-layer position remains unfalsified rather than supported, which is a different state that gets spoken about identically.
Our own result being unpopular. The null we published is not a claim that schema does nothing. It is a claim that the effect, if present, is smaller than the instruments currently pointed at it can resolve. Those are different sentences and we would rather be quoted on the second one.
What our own instrument still cannot do
Our published measurement establishes two things and settles none of the questions in this post. It shows that identical queries return different citations, at 7 of 12 on our own set. It shows what no effect looks like, across 20,000 permutation splits. It cannot show whether schema causes citations, because its detection floor is 51.1 percentage points.
We are stating this in its own section rather than in a footnote because the alternative is the shape this whole post objects to. A methodology piece that ends by implying its author has the answer is doing the same trick as a case study that ends by implying causation, just with better vocabulary.
The specific limits, so a reader can weigh them:
- Small query set. 12 queries is enough to demonstrate instability and nowhere near enough to estimate an effect. The flip rate itself is measured with wide uncertainty.
- One model, one day. We deliberately compressed the window to remove model-drift as a confound, which costs generalisability. We do not know how the flip rate behaves across models or across weeks.
- No treatment arm at all. Both published runs are untreated. We have measured the noise, not a signal against it.
- No customer engagements yet. Nothing in this post is a client result, and there is no engagement whose numbers we are withholding. There are none to withhold.
What we would need to move from here is straightforward and expensive: a pre-registered query set of 40 or more, 40 repeats, a matched held-out control arm, and a treatment applied to pages that are not currently cited. That design reaches a 9.9-point floor, which is honest enough to publish and still not sensitive enough to detect the small effects that are most plausible. That is the state of the art, and it is worse than the state of the marketing.
Provenance notes on this post's own sourcing
Three things in the reporting above are weaker than the rest and we would rather name them than let a reader discover them.
The DUCKYEA experiment has no primary link here. The fictional-company test described in the tokenisation section is widely reported and we did not resolve its original post to a URL we could read on 2026-09-09, so it is described rather than cited. The tokenisation argument it supports does have a primary source, linked, and that is the load-bearing half.
The April 2025 Google statement is second-hand. The line about structured data giving an advantage in AI Overviews reaches us through Search Engine Land's reporting rather than through a Google surface we opened. We treat it as accurately reported and not as first-party, which is why the section above compares it to the December 2025 documentation on requirement rather than declaring one of them wrong.
The Microsoft position is unlocated. A March 2025 Bing statement that schema helps large language models understand content is cited in the same piece. We looked for the original and did not find it. It stands at second-hand authority only and we have not built anything on it.
None of the three changes a conclusion in this post. Saying so is cheaper than having somebody find them.
The blunt answer
Ship schema markup because it earns rich results and describes your page accurately, not because it moves AI citations. The largest controlled test found no uplift, our own extraction test found no engine reading it on a direct fetch, and nobody selling you a citation lift from structured data has published a control arm.
The long version, one clause at a time:
- Google says no special structured data is needed for AI features, on its own live documentation, updated 2025-12-10.
- The largest controlled test, 1,885 treated pages against 4,000 matched controls, found no uplift on any platform, and its own authors say they cannot tell whether schema did a tiny bit of good or nothing at all.
- Independent extraction testing found that a fact placed only inside JSON-LD was returned by zero of five AI systems during a direct fetch. Everything they read was visible on the page.
- The correlational evidence is real and consistently positive, and it is confounded by the fact that competent sites do ten things right at once.
- The retrieval-layer argument for a schema effect is unfalsified, which is not the same as supported, and the test that would settle it has not been run by anyone including us.
- Your own before-and-after cannot tell you the answer, because identical queries flip on their own more than half the time, and a raw plus 43 percent in that dataset was plus 2.4 percent once the control group was subtracted.
The reason to ship schema anyway is that it does other things, and those other things are measured, established and older than this argument. Rich results, entity resolution, merchant surfaces. Ship it for those, keep it matching your visible text, and put your measurement budget somewhere with a larger prior.
If you want the noise band on your own queries before you spend anything, that is the first thing we run and we will run it for four queries at no cost. It is also the number that decides whether any test you commission afterwards can mean anything, which is why we do it before anything else rather than as an upsell. The rest of what that engagement involves is on answer engine optimization, and what it costs is on pricing.
One last note on this post's own markup, since it would be strange not to. It ships Article, FAQPage, BreadcrumbList and Dataset. Every claim in the JSON-LD has a visible counterpart on the page. We do not expect any of it to earn a citation, and if this post gets cited we will not be able to tell you it was the schema. That is the whole argument, applied to itself.
Further reading in this cluster: what answer engine optimization actually means for the definitional groundwork, and original research on AI citations for everything else we have published with its numbers attached.
Sources
Every number above, and where it came from. A figure without a row here is one we should not have printed.
- Ahrefs, "We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved."
- Louise Linehan and Xibeijia Guan, published 2026-05-11, reviewed by Ryan Law. 1,885 pages that added JSON-LD between August 2025 and March 2026, each matched to 3 control URLs from different domains at similar pre-period citation levels, analysed with a matched difference-in-differences test. Google AI Overviews -4.6 percent, AI Mode +2.4 percent, ChatGPT +2.2 percent. Read in full, 2026-09-09.
- searchVIU, "Schema Markup and AI in 2025, What ChatGPT, Claude, Perplexity and Gemini Really See"
- Michael at searchVIU, published 2025-12-02, tests run October 2025. Eight price-extraction tests on one purpose-built page across five AI systems. The price that existed only inside JSON-LD was found by zero of the five. Read in full, 2026-09-09.
- Google Search Central, "AI features and your website"
- Last updated 2025-12-10 UTC. Verbatim: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." The same page also recommends "making sure your structured data matches the visible text on the page." Both lines read directly off the live document, 2026-09-09.
- Kumar and Palkhouski, "AI Answer Engine Citation Behavior: An Empirical Analysis of the GEO-16 Framework"
- arXiv preprint 2509.10762, dated 2025-09-13. 70 product-intent prompts, 1,702 citations across Brave Summary, Google AI Overviews and Perplexity, 1,100 unique URLs audited. Structured Data is among the pillars showing the strongest association with citation. The abstract states plainly that the study is observational. Abstract read directly, 2026-09-09.
- Otterly.ai, "The real impact of schema markup on AI search"
- Published 2026-03-23. Schema shipped 2025-12-07 and tracked to 2026-03-07, 319 monitored prompts, 7 platforms, 5 schema types. Reports that 6 of 7 AI search platforms could not fetch or correctly interpret schema when asked directly, and attributes its AI Overview movement to traditional-search mechanics because competitors who changed no schema moved in parallel.
- Search Engine Land, "Schema markup and AI search, no hype"
- Published 2026-03-25. Frames schema as infrastructure rather than a lever, cites a December 2024 Search/Atlas study that found no correlation between schema coverage and citation rates, and states that there are no peer-reviewed studies on schema's impact on AI search visibility.
- Search Engine Roundtable, "Structured data and schema for AI search visibility"
- Reports Mark Williams-Cook's tokenisation demonstration, in which the token boundaries a model applies to a JSON-LD block split '"@type" - "Organization"' into separate tokens for "type" and "Organization", making the markup indistinguishable from the same two words written in a sentence.
- Mark Williams-Cook on LinkedIn, the original tokenisation post
- The primary post behind the tokenisation argument, a visual explanation of why a JSON-LD block does not survive tokenisation as structure. Published September 2025.
- Evertune, "Schema vs. No Schema, Does Structured Data Matter for AI Search?"
- Published 2025-11-06, roughly 1,033 words. Cited here as the gap rather than as evidence. The title promises the comparison and the body does not run it, citing unnamed "recent experiments" with no design, sample size or numbers. It does concede that large language models may strip schema markup during tokenization.
- WordStream, "Schema Markup for AI, Types, Benefits, and How to Set It Up"
- Roughly 2,805 words, the longest credible article on the live page one for this query as measured 2026-09-09. A type-by-type setup walkthrough plus a validation workflow, with no experiment, no control group and no numbers of its own.
- SEOptimer, "Schema Markup for AI Search: Complete Guide"
- Roughly 2,177 words. Competent implementation guide whose five quoted statistics are all borrowed and none of which is about AI-citation causality, including the "72 percent of first-page results use schema" figure and a 40 percent CTR claim sourced from a schema vendor's own marketing.
- Wix Studio AI Search Lab, "Schema markup in AI search"
- Vendor-authored. The numbers it carries are uncontrolled single-site before-and-after case studies with no disclosed methodology, including a 19.72 percent increase in AI Overview visibility measured on the vendor's own site.
- TheHoth, "Schema markup for AI"
- Roughly 2,178 words, a new entrant to page one since the 2026-09-02 capture and the closest title-level competitor to this post's angle. Verified 2026-09-09 by reading it, 12 headings and zero hits for methodology, sample size, experiment or baseline.
- Schema App, "The future of search, AI, machine learning and schema markup"
- Vendor-authored and the oldest page on the SERP, roughly 1,728 words. Sourced from conference keynotes and quoted chatbot responses, with no study and no sample size.
- Google Search Central, FAQPage structured data documentation
- The current eligibility rules for FAQ rich results, which Google narrowed in August 2023 to a small set of well-known authoritative government and health sites. The markup remains valid; the rich result is what went away for nearly everybody.
- Google Search Central blog, "Changes to HowTo and FAQ rich results"
- The August 2023 announcement that removed FAQ rich results for most sites and HowTo rich results entirely on desktop. The primary source for why FAQPage markup is now an extractability decision rather than a rich-result decision.
- schema.org, the Dataset type
- The vocabulary definition for Dataset, the type this post ships on its own measurement runs and the one that no page currently ranking for this query uses.
- Ahrefs, "What is Query Fan-Out?"
- Despina Gavoyannis, 2026-03-02. Background on the fan-out step that sits between a user's question and the pages an AI answer cites, which is the layer where a retrieval-side schema effect would have to live if one existed.
- r/DigitalMarketing, "Google published its official guide on getting cited by AI"
- 124 upvotes, 79 comments, posted 2026-06-08. A practitioner in the AI-visibility industry arguing against their own industry's pitch. Cited as market voice, not as a statistic. Counts read live via our own Reddit data tooling, 2026-09-09.
- r/TechSEO, "Ahrefs Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved."
- 17 upvotes and 60 comments, posted 2026-05-11, upvote ratio 0.66. The comment-to-upvote ratio and the split ratio are the tell that this is contested rather than settled. Counts read live, 2026-09-09.
- r/aeo, "We tested 50 sites with GTM-injected structured data"
- A practitioner audit reporting that 34 of 50 sites injected JSON-LD through Google Tag Manager or other client-side JavaScript, that 0 percent of GTM-only sites had schema detected by AI crawlers against 92 percent of server-rendered ones, and that some were paying for the invisible markup. Cited as a field report, not as a controlled study.
- Aaron Haynes, Partner and CEO at Loganix, on X
- The strongest public statement of the pro-schema position and the one this post engages with directly, including the reconciliation that the mechanism is not the transformer's comprehension but the retrieval pipeline in front of it. Read live via our own X data tooling, 2026-09-09.
- Our own step-zero measurement run
- 12 money queries times 5 identical repeats, one model, one day. 60 of 60 calls succeeded and 7 of the 12 queries flipped outcome between identical repeats with no intervention applied. Published with its numbers on this site.
- Our own permutation test on the same query set
- 20,000 random splits of the same query set with no intervention applied. The null difference is centred on zero, mean plus or minus 0.0016, standard deviation 0.215. The same post states the design's minimum detectable lift at 51.1 percentage points, roughly four times underpowered for the effects the market claims.
Questions this answers
- Does schema markup help with AI search?
- There is no public evidence that it causes AI citations directly. Google's own AI features documentation says no special schema.org structured data is needed to appear in AI Overviews or AI Mode, and the largest controlled test found no uplift across 1,885 pages that added JSON-LD.
- Does Google require structured data to appear in AI Overviews?
- No. Google's AI features page, last updated 2025-12-10, states that you do not need to create new machine readable files, AI text files or markup to appear in these features, and that there is no special schema.org structured data you need to add.
- Can an AI assistant actually read the JSON-LD on a page?
- Not during a direct fetch. In eight extraction tests across five AI systems, a price that existed only inside JSON-LD was returned by none of them. The same systems read visible HTML, visible Microdata and visible RDFa. Search indexes do extract structured data, which is a different path.
- Which schema types are most useful for AI search?
- The useful list is the traditional search list. Organization, Article, Product, BreadcrumbList, LocalBusiness, and Dataset for research pages. No type has been shown to move AI citations specifically. Choose types that describe what is visibly on the page, because a markup claim with no visible counterpart is a liability.
- Is FAQPage schema still worth adding for AI answers?
- For rich results, mostly not. Google narrowed FAQ rich results in August 2023 to a small set of authoritative sites. For machine extractability it can still earn its keep, because a question and answer block is easy for any parser to lift. Ship it for that reason and expect nothing else.
- How do you isolate schema as a variable in an AI citation test?
- Hold everything else still. Matched control pages that get no schema, a query set fixed before you look, every query repeated enough times to see the noise band, randomised assignment of which pages get treated, and no other edits inside the window. Without all five you have a coincidence.
- How many repeats does an AI citation test need before the result means anything?
- Enough that the noise band is visible before the treatment ships. In our own step-zero run, five repeats of twelve queries on one model in one day was enough to show that seven of the twelve flipped on their own. Five repeats measures the problem. It does not solve it.
- What is a permutation test and why does it matter for AI visibility claims?
- You shuffle your own data into random groups thousands of times with no treatment applied, and record how big a difference appears anyway. That is your null distribution. Ours, across 20,000 splits, centres on zero with a standard deviation of 0.215. Any real result has to beat it.
- Why do identical AI queries return different citations?
- Because the systems are not deterministic and their retrieval step re-runs each time. In our step-zero run of 12 queries repeated 5 times each on one model on one day, 60 of 60 calls succeeded and 7 of the 12 queries flipped outcome with nothing changed at all.
- How small a citation lift can a small test actually detect?
- Usually far larger than the lift being claimed. A 12-query, 5-repeat design has a minimum detectable lift of 51.1 percentage points, roughly four times underpowered. Moving to 40 queries and 40 repeats brings it to 9.9 points, which is better and still not small.
- Does schema markup affect ChatGPT citations differently from Google AI Overviews?
- The controlled test reports different numbers per platform, plus 2.2 percent for ChatGPT and minus 4.6 percent for AI Overviews, but only the AI Overviews figure was statistically distinguishable from zero and its authors decline to attribute it to schema. Read that as no reliable difference.
- Is a before and after schema rollout enough evidence that schema caused a lift?
- No. Without a control group you cannot separate your change from an algorithm update, a recrawl, a freshness effect or ordinary volatility. In the controlled test, a raw before-and-after on AI Mode read plus 43 percent and collapsed to plus 2.4 percent once matched controls were applied.
Keep reading
AI citations
LLM SEO, what it actually is and what the work looks like
Google says its AI features run on the same ranking systems as Search, so the tactics are familiar. What changed is how you tell whether the work landed.
46 min read
AI citations
Reddit is cited by AI. That is not a reason to buy upvotes.
Published Reddit citation shares run from 2% to 46.7%, and one study puts Reddit at 67.8% of every URL ChatGPT retrieves and then declines to cite.
23 min read
AI citations
Does llms.txt actually move AI citations, and how would you tell
Every page ranking for this term explains what llms.txt is. None answers whether it changes anything. What 40 sites deploy, and the test that would tell.
43 min read