Does Schema Markup Actually Get You Cited by AI? Here's What the Evidence Shows
In October 2025, a researcher at searchVIU built a fake product page and hid the price — €8.99 — inside JSON-LD schema markup only. Not in the visible text, not in an image, nowhere a human eye would ever see it. Then he asked five AI systems, one at a time, what the product cost. Claude found it zero times out of eight tries. Perplexity found it once. Google AI Mode found it twice. ChatGPT got it three times. Gemini, the only one of the five that renders JavaScript during a live fetch, got it four times out of eight — and even that was a coin flip. You can read the full write-up on searchVIU's site if you want the exact test matrix. The headline finding is blunt: when these systems fetch a page in real time to answer a question, none of them read the schema. They read the visible HTML and nothing else.
That's an awkward fact for an industry that's spent the last two years telling businesses schema markup is the secret handshake that gets you cited by ChatGPT. It's also not the whole story — and figuring out where the honest line sits between "this genuinely matters" and "this is theater" is the actual point of this piece. Because there are two separate questions people keep collapsing into one: does schema help Google rank you higher, and does schema help AI systems cite you. Google has answered the first question directly, repeatedly, on the record. The second question has real evidence too, and it points somewhere more interesting than either the hype or the dismissal.
Google Has Said the Quiet Part Out Loud, More Than Once
Start with what's actually settled, because it's more settled than most agencies let on. In April 2025, Google's John Mueller was asked on Bluesky whether structured data affects ranking, and he didn't hedge: "Structured data won't make your site rank better," he wrote, adding that its job is displaying the specific search features documented on Google's own developer site — nothing more. That wasn't a new position. Barry Schwartz's write-up on Search Engine Roundtable traces the same claim back to 2018, when Mueller said there's no generic ranking boost from markup, though it "can make it easier to understand what the page is about."
Google's own developer documentation backs this up without any spin. The introduction to structured data in Search Central frames the entire benefit around rich results — the star ratings, FAQ dropdowns, and recipe cards that make a listing look more clickable — and the case studies it cites (Rotten Tomatoes, Food Network) all measure click-through rate and engagement, never ranking position. And when Google published its own guidance on AI features and your website, it went a step further than the ranking disclaimer: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." Read that twice. That's Google explicitly telling you schema isn't a prerequisite for showing up in AI Overviews either.
So if you came here hoping schema markup was a ranking hack or a guaranteed AI Overview ticket, that's settled, and it's settled by Google itself, not by us. But "not a ranking factor" was never really the question that mattered for AI citation. Ranking and citation are governed by different mechanics now — Semrush has documented ChatGPT regularly citing pages sitting at position 21 or lower in Google's own results — so the fact that schema doesn't move one doesn't tell you anything about the other. That's the actual question worth digging into.
The Correlation Looks Strong — Almost Too Strong
Here's where it gets genuinely interesting. Ahrefs published a study in May 2026, by Louise Linehan with Xibeijia Guan, that opened with a striking number: pages cited by AI systems were almost three times more likely to carry JSON-LD schema than pages that weren't cited. Separately, an academic paper out of UC Berkeley — "AI Answer Engine Citation Behavior," building a 16-pillar scoring framework called GEO-16 across 1,100 audited URLs and 1,702 citations pulled from Brave, Google AI Overviews, and Perplexity — found the Structured Data pillar correlated with citation likelihood at r = 0.63 (p < 0.001), one of the strongest single correlations in their whole model. You can read the full paper on arXiv. On the surface, that's a compelling case: businesses with clean schema get cited more.
But both of those are correlational studies, and both sets of authors are honest enough to say so out loud, which is rarer than it should be in this field. The Berkeley researchers write plainly in their limitations section that their "observational design may suffer from unobserved confounding" — meaning the businesses with tidy JSON-LD markup are probably also the businesses with better content, cleaner sites, and more established brands generally, and any of those things could be doing the actual work. Correlation this strong is a real signal worth noting. It's not proof of cause. Anyone selling you the r = 0.63 number as if it settles the argument is skipping the part of the paper where the authors tell you not to do that.
Then Someone Actually Tested It — and the Needle Barely Moved
This is the part that makes the Ahrefs research worth reading in full rather than just quoting the headline stat. Correlational data can't tell you what happens if you add schema to a page that doesn't have it. So Ahrefs ran the experiment: they tracked 1,885 real pages that added JSON-LD schema between August 2025 and March 2026, matched against roughly 4,000 control pages, and measured AI citation counts 30 days before and after, using four separate statistical methods (t-tests, difference-in-differences, an event study, and a sensitivity check) to make sure one shaky method wasn't driving the result.
The outcome: Google AI Overviews citations dropped 4.6% on treated pages — a small, statistically real decline, working out to roughly 12 fewer daily citations per page on average. Google AI Mode moved +2.4% and ChatGPT moved +2.2%, both statistically indistinguishable from zero — noise, not signal. The authors' own summary is refreshingly unhedged: "Adding schema produced no major uplift in citations on any platform." And this lines up exactly with the searchVIU finding from the opening of this piece — the same Ahrefs article cites that experiment directly, noting that during live retrieval, "every system extracted only visible HTML content. JSON-LD, hidden Microdata, and hidden RDFa were all ignored." If the systems doing the citing don't read the markup when they visit your page, it stands to reason that adding more of it wouldn't move the needle on pages they're already visiting and already citing.
There's a real limitation worth being upfront about, and Ahrefs names it themselves: every single page in their sample already had 100+ AI Overview citations before the schema was added. This is a study of whether schema helps pages that are already succeeding succeed more. It says nothing about whether schema helps a page get into the running in the first place — get crawled, get its entity understood, get considered as a candidate source at all. That's a different, earlier stage of the pipeline, and nobody has run a clean causal test on it yet.
"Infrastructure, Not a Lever" — Where the Actual Disagreement Lives
That gap is exactly what Gianluca Fiorelli, a well-known international SEO consultant, picked apart in a direct response to the Ahrefs study. His piece, "The Ahrefs Schema study is right. And it's testing the wrong thing," doesn't dispute a single number Ahrefs published. "Ahrefs is right," he writes. "Within the scope of what they tested, the data is credible, and the conclusion holds." His argument is about scope: every page in the dataset was already in what he calls the AI's "consideration set" — already crawled, already surfaced, already known. Fiorelli's framing is the sharpest sentence in the whole debate: schema's real work "happens earlier in the pipeline, at the indexing and entity resolution stage" — not as something that nudges an already-established page higher, but as the plumbing that determines whether your business gets correctly identified as a distinct, coherent entity in the first place. His phrase for it is schema as "infrastructure, not a lever" — closer to registering your company than to running an ad.
You can see the same instinct, minus the academic rigor, in a real conversation happening on Hacker News. On a Show HN thread for a free AI-discoverability audit tool, the builder, posting as lr001328, wrote plainly: "Structured data matters most. Sites with proper JSON-LD schema see measurably higher AI citation rates. Microsoft has confirmed schema markup helps their LLMs." Worth noting: the same person, in the same post, was openly skeptical about a different tactic — llms.txt files — writing "we should be honest: no major AI platform has publicly confirmed they read it, and statistical analysis shows no correlation with citation rates." That's a person willing to call out hype when they see it, and they still rank schema as the thing that "matters most." On a separate thread about a small SaaS getting picked up in Google's AI Overview, a commenter using the handle 13pixels made a more hedged version of the same point: clear Schema markup for a product or service "seems to help LLMs 'understand' the entity better, even if they don't rank the page itself for the query. It's like giving them a cheat sheet." And on a broader thread about AI reshaping SEO, a commenter posting as kiji-seo summed up the state of the industry about as honestly as anyone: "there is a real split within the SEO community right now, those who understand that the AI is causing a major shift in SEO focus and those who believe that it is still 2015" — and named structural work like schema as exactly the kind of thing the first group has started prioritizing.
None of that is proof. It's practitioners describing what they're seeing, hedged the way honest people hedge when they don't have a controlled experiment to point to. Which, again, is the whole shape of the evidence here: real correlation, real caution about causation, and real disagreement about which stage of the pipeline actually benefits.
What This Actually Means If You're Deciding Whether to Bother
Put all of it together and a fairly specific, non-generic picture holds up. Schema markup will not move your Google ranking — Google has said this plainly enough times that arguing otherwise means arguing with the company that runs the ranking system. It will not, on its own, cause an AI Overview citation spike on a page that's already succeeding — Ahrefs tested that directly and the number came back essentially flat. AI systems don't even read your JSON-LD when they visit your page live — searchVIU showed that with real, replicable tests across five platforms. If someone tells you adding schema this quarter is going to be the thing that gets you cited by ChatGPT next quarter, the actual data says otherwise.
What schema does appear to do — supported by a real correlation, an honest academic caveat about confounding, and a specific, named critique of the Ahrefs study's scope — is function earlier and quieter than a ranking lever. It's part of how a system resolves who you actually are: which business this is, what it does, where it operates, whether the "ABC Dental" mentioned on three different pages is the same ABC Dental or three different ones. That's disambiguation work, not persuasion work, and disambiguation matters most exactly when there's ambiguity to resolve — a common business name, multiple locations, a brand that shares a name with something else entirely. A business with a genuinely distinctive name and one clear location has less ambiguity to clear up in the first place, which is part of why the effect is so hard to isolate in aggregate data: the value is real but concentrated, not evenly spread across every business that implements it.
That's also, not coincidentally, close to what Google itself says structured data is good for — not ranking, but making sure Google (and, it now looks like, other systems downstream of similar signals) can confidently parse what your business actually is. It's the same logic behind why we scope our own Schema Markup Services around correctness and validation rather than promising a citation bump: the evidence supports schema as groundwork that removes ambiguity, not as a lever you pull for more visibility. If you want to check where your own markup currently stands before paying anyone to touch it, our free schema validator tool will tell you in a couple of minutes whether what's already on your site is even valid — which, per the searchVIU test, matters less for how AI systems read your page live and more for whether Google can build an accurate picture of your entity in the first place.
So Should You Actually Bother?
If you're weighing schema work against genuinely more urgent problems — a site that's slow, content that doesn't actually answer the question it's targeting, an FAQ page that doesn't exist yet — do those first. The evidence doesn't support schema as a shortcut past better content, and nobody credible in any of the sources above claims otherwise. But if your fundamentals are already solid and you're specifically dealing with ambiguity — a multi-location business, a name shared with other companies, a brand where "who exactly is this" isn't obvious from prose alone — implementing FAQPage, LocalBusiness, Organization, and Review schema correctly is cheap, low-risk, and grounded in a real (if indirect) mechanism, even though it won't show up as a ranking bump or an overnight citation spike in your analytics next week. Treat it the way Fiorelli suggests: as infrastructure you build once and validate properly, not a lever you pull and expect to see move.
Frequently asked questions
Ghulam Mustafa is the founder of AI Rankings and CEO of a digital marketing agency based in Abu Dhabi, UAE. His career sits at the intersection of full-stack development and search — building on Flask, Django, WordPress, and JavaScript while running SEO, AEO, and GEO campaigns for clients across the region. AI Rankings grew out of that work: a platform for tracking how brands actually show up in AI-generated answers, built on the principle that every number it reports has to be real and verifiable, never estimated or simulated. He writes about AI search visibility, technical SEO, and the shift from ranking on Google to being cited by AI.
View profile →See where you actually stand right now.
Free, live check. Real evidence, not an estimate.