What the GEO Research Actually Tested

Ghulam Mustafa
Ghulam Mustafa — Founder
· September 9, 2026

Nine techniques. One benchmark of 10,000 real queries. And one of the nine — keyword stuffing, the oldest trick in the SEO book — made content less likely to get cited, not more. That's from the actual paper that coined the term "GEO," and it's the kind of detail that gets dropped from almost every agency blog post that cites the headline "up to 40%" number without ever opening the PDF. If you're technical enough to want the real methodology instead of the marketing summary, this is that piece.

The paper is "GEO: Generative Engine Optimization" by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande — Princeton and Georgia Tech researchers, first posted to arXiv in November 2023, revised through mid-2024, and formally presented at ACM SIGKDD 2024. It's the study almost every "GEO service" now traces its legitimacy back to. Most of what gets built on top of it, though, either oversimplifies what it tested or quietly ignores the parts that don't support a clean sales pitch. Worth fixing that.

What GEO-Bench Actually Is

The researchers didn't just run a few example queries and call it a study. They built a benchmark — GEO-bench — out of 10,000 queries pulled from nine different sources: real user searches from Bing and Google via MS MARCO and ORCAS-1, Natural Questions, Oxford's AllSouls essay questions, the reasoning-heavy LIMA dataset, debate prompts, trending queries pulled from Perplexity.ai's own Discover feed, Reddit's ELI5 threads, and a batch of GPT-4-generated queries built to cover specific domains and intents. Split 8,000/1,000/1,000 for training, validation, and test.

That mix matters more than it sounds like it should. Eighty percent of the queries are informational, 10% transactional, 10% navigational — tagged across 25 domains spanning Arts, Health, Games, and more, and cross-labeled across seven different categorization schemes. It's a genuinely broad test set, not a cherry-picked pile of queries where a specific technique was already known to work. That's the part that makes the "up to 40%" number defensible rather than a stat someone pulled from a single favorable case.

The Nine Techniques, and What Actually Moved the Needle

Here's the part that matters if you're deciding what to actually spend time doing. Against a no-optimization baseline scoring 19.3% on the paper's core metric — Position-Adjusted Word Count, essentially how much of the AI's answer draws from your content and how prominently — the nine techniques landed in three tiers:

  • Quotation addition (adding direct quotes from credible sources): 27.2%, a +40.9% relative lift — the single best-performing technique in the whole study.
  • Fluency optimization (rewriting for cleaner, more natural prose): 24.7%, +28.0%.
  • Cite sources (working in citations from credible outside sources): 24.6%, +27.5%.
  • Statistics addition (swapping vague claims for actual numbers): 25.2%, +30.6%.
  • Technical terms (domain-specific vocabulary): 22.7%, +17.6%.
  • Easy-to-understand (simplifying language): 22.0%, +13.9%.
  • Authoritative (more persuasive, confident tone): 21.3%, +10.3%.
  • Unique words (unusual terminology): 20.5%, a marginal +6.2%.
  • Keyword stuffing (more query-relevant keywords, the classic SEO move): 17.7% — a -8.3% drop from baseline. The only technique of the nine that actively hurt.

That last line is the one worth sitting with. The single most reliable habit of two decades of traditional SEO — cram in the keyword — actively backfired in this benchmark. Meanwhile the four techniques that clustered at the top all share something in common: they add real, checkable specificity (a quote, a stat, a citation, cleaner writing) rather than trying to game a ranking signal. That's a genuinely different game than the one most agencies were trained to play, and it's worth remembering the next time someone tries to sell you an "AEO content calendar" that's really just more keyword-optimized blog posts wearing a new label.

The Part Most Summaries Skip: How the Test Actually Worked

Here's where the paper gets narrower than the marketing built on top of it, and it's a distinction worth being precise about. The researchers used what they call a two-step setup: for each query, a real Google search retrieves the top five candidate sources, and then GPT-3.5-turbo (temperature 0.7, five sampled responses per query) generates the actual answer, drawing on whichever of those five sources it chooses to cite. The nine techniques were then applied to rewrite one of those five sources and the whole thing was re-run to see if the rewritten version pulled more weight in the output.

Notice what that setup assumes: your content is already one of the five sources the search step retrieved. The paper is testing what happens once you're in the room, not what gets you into the room in the first place — and the authors are upfront about why, citing "context length limitations and quadratic scaling cost based on the context size of transformer models" as the reason only five sources get fetched at all. That's a real, sensible constraint for running a controlled benchmark. It also means the strict academic definition of GEO is about synthesis and citation quality among sources already retrieved — not about crawlability, indexing, or whether a generative engine's retrieval step surfaces your page as a candidate to begin with. That earlier, upstream layer is a separate mechanism, one we've covered in more depth on our own Generative Engine Optimization page — this piece is specifically about what the original research tested and how well that's held up since.

The authors also ran one real-world check outside the controlled benchmark: applying GEO techniques against Perplexity.ai as a live, commercially deployed generative engine, where they report visibility improvements of up to 37% — close enough to the benchmark number to suggest the lab result wasn't purely an artifact of the test setup, though it's a single spot-check against one platform rather than a second full benchmark.

What Happened When Other Researchers Actually Tried to Check This

A study getting cited constantly isn't the same as a study holding up under scrutiny, so it's worth asking what's happened since 2024. The most thorough answer so far comes from a July 2026 paper, "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)" by Olivier Martinez, which reviewed 45 separate studies published in the field since the original GEO paper came out.

Its conclusion isn't a debunking, but it is a serious tightening of the claim. In the survey's own words, the original paper's "widely cited gains are valid within its experimental setting but conditional on a source already being present in a fixed context; they establish neither organic discoverability nor durable traffic effects." In plain terms: the 40% number is real, but it only tells you what happens once you've already made the shortlist — exactly the gap the two-step setup above leaves open. The survey goes further, reporting that across the 45 studies it reviewed, generic optimization heuristics "transfer poorly" between domains and platforms, competitive pressure from other sites doing the same optimization erodes individual gains over time, and — this is the uncomfortable one — some citation-oriented rewrites actually impair how well a page gets retrieved in the first place, trading one metric for another. Commercial visibility-tracking tools, the survey adds, show low overlap between what different tools report as "cited" and substantial run-to-run variability even measuring the exact same query twice.

"no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior"

Olivier Martinez, "Optimizing Visibility in Generative Engines," July 2026

That's a genuinely important caveat for anyone quoting "GEO boosts visibility by 40%" as if it's a settled, durable fact rather than a specific, replicable result inside a specific experimental frame. The original researchers never claimed otherwise — the survey's real complaint is with everyone downstream who stripped the caveats off on the way to a sales page.

A Different Team Found Similar Patterns, With a Twist

Not everything since 2024 is critique. A Carnegie Mellon team — Yujiang Wu, Shanshan Zhong, Yubin Kim, and Chenyan Xiong — published "What Generative Search Engines Like and How to Optimize Web Content Cooperatively" in October 2025, building a framework called AutoGEO that automatically discovers what a given generative engine actually rewards and rewrites content accordingly. Their system landed on more than 20 preference rules, and the broad strokes echo the original Princeton findings closely: comprehensive coverage, credible source attribution, specific data and statistics, clear structure, and a neutral rather than promotional tone all correlated with higher citation rates. Using engine-specific rule sets rather than one generic checklist, AutoGEO produced an average 35.99% improvement in visibility — genuinely close to the original paper's 40% ceiling, from an independent team two years later.

The twist is the domain-specific part, which the original paper flagged as a limitation without fully quantifying. AutoGEO found that e-commerce-oriented engines rewarded step-by-step guides and concrete recommendations, while research-focused contexts rewarded explanatory depth and balanced framing of competing views. Same underlying mechanism, meaningfully different content depending on what you're actually optimizing for — which is a real problem for any "one-size-fits-all AI content checklist" an agency hands you without asking what kind of query your business actually gets asked.

What Technical People Are Actually Arguing About

The cleanest real-world debate about this research showed up recently in a Launch HN thread, not a marketing blog. In "Launch HN: Sitefire (YC W26) — Automating actions to improve AI visibility," founder vincko described a platform that monitors AI-generated "fan-out queries" across ChatGPT, Gemini, and Google AI Mode and pushes content changes straight into a client's CMS — reporting that for one client, AI-optimized articles pushed bot requests from roughly 200 a day to 570 a day within ten days.

The comments got specific fast. A commenter using the handle onecommit asked the question that cuts right to the survey's core critique before the survey existed: "How do models deal with assessing the quality of content and its accuracy/veracity when recommending products currently? What do the providers do to avoid a situation where more content === more traffic? Would love to see links to relevant research on this, if you have them." vincko answered by linking the original Princeton paper directly and summarizing its approach: "This paper tried different strategies to improve performance for this last step... Adding statistics, sources, original data are all strategies that we apply. In classic SEO, creating more and more content leads to 'cannibalization'. Generally this hurts performance of all overlapping content so much that it is not worth it."

Another commenter, marzapower, pushed on methodology in a way that lines up almost exactly with the academic survey's later findings — arguing that prompt-monitoring itself is shaky ground: "personalization, location variance, account history all introduce noise that's hard to control," and proposing instead "page-level structural analysis: instead of asking ChatGPT 'do you cite this site?', you analyze the page directly for the signals that predict citation — source density, answer structure, fluency, statistics... based on the Princeton KDD research." And a more skeptical commenter, ceejayoz, questioned the whole premise on incentive grounds: the customer paying for AI-visibility work is the business, not the person asking the AI a question, "who'll tend to accept whatever's the top search result even if it's deeply wrong or complete slop" — a fair echo of the same worry that's dogged SEO for two decades, just aimed at a newer target.

Nobody in that thread thinks the research is wrong. What they're actually arguing about is whether the measurement layer sitting on top of it — the tools and dashboards claiming to track "AI visibility" in real time — is solid ground to build decisions on, or noisy enough that the honest answer is "probably directionally right, don't trust the exact number."

So What Should You Actually Take From This?

Strip away the caveats and a few genuinely actionable conclusions survive contact with the follow-up research. Specificity beats vagueness, consistently, across every study cited here — quotes, statistics, and citations all outperformed generic persuasive writing by a wide margin, and that finding has now been independently reproduced by a different team on different models eighteen months later. Keyword stuffing doesn't just fail to help; it measurably hurts, which should finally put to rest the instinct to treat GEO as SEO with a new coat of paint. And the gains the original paper measured are real but conditional — they describe what happens once your content has already been retrieved as a candidate, not whether it gets retrieved at all, which is a separate, upstream problem involving entity consistency, structured data, and crawlability rather than clever rewriting.

The honest caveat, and it's a real one: nobody has yet published a study showing these techniques produce a stable, long-term, cross-platform lift in actual organic discoverability — the July 2026 survey is explicit about that gap, and it's not a small one. What exists instead is strong, repeated evidence for the narrower claim: specific, well-sourced, clearly structured writing wins the synthesis step once you're in the candidate pool, and generic or keyword-stuffed writing loses it. That's a real, useful, evidence-backed thing to build a content strategy around — it's just not the whole picture, and anyone telling you it is hasn't read past the abstract.

If you want to see where your own content actually stands against these specific mechanics before rewriting anything, the free AEO Score tool is a faster starting point than guessing from a paper's headline number.

Frequently asked questions

What is GEO-bench, the benchmark used in the original GEO paper?
GEO-bench is a 10,000-query benchmark built by the Princeton/Georgia Tech researchers from nine sources, including MS MARCO, ORCAS-1, Natural Questions, Oxford's AllSouls essay questions, LIMA, debate prompts, Perplexity.ai Discover queries, Reddit's ELI5, and GPT-4-generated queries. It's split 8,000/1,000/1,000 for training, validation, and test, covering 25 domains and mostly informational queries (80%).
Which content technique performed best in the study, and by how much?
Quotation addition (adding direct quotes from credible sources) performed best, improving visibility by 40.9% relative to a no-optimization baseline. Statistics addition (+30.6%), fluency optimization (+28.0%), and cite sources (+27.5%) followed closely behind.
Did any technique actually hurt visibility?
Yes. Keyword stuffing was the only one of the nine techniques tested that decreased performance, dropping visibility by 8.3% below the baseline. Unique-word insertion barely helped at all, at +6.2%.
Does the original GEO paper measure whether AI systems find your content in the first place?
No. The paper uses a two-step setup where Google search first retrieves the top 5 candidate sources, and the tested techniques only affect how well a source performs once it's already among those five. It measures synthesis and citation quality among already-retrieved sources, not discoverability or crawlability.
Has anyone checked whether the original GEO findings hold up since 2024?
A July 2026 critical survey by Olivier Martinez reviewed 45 studies published since the original paper and found its gains are real but conditional on a source already being in a fixed context, don't establish organic discoverability, and can erode under competitive pressure. A separate October 2025 Carnegie Mellon study (AutoGEO) independently found similar techniques produced a 35.99% average visibility improvement.
Is GEO the same thing as AEO?
Not exactly. GEO, in the strict academic sense from this paper, is narrower: it's specifically about optimizing content so a generative engine synthesizes and cites it well once retrieved. AEO, as most agencies use the term today, is the broader practice covering retrieval, entity consistency, and structured data too, not just the writing itself.
What are people debating about this research in public technical forums?
On a Hacker News Launch thread for the AI-visibility startup Sitefire, commenters debated whether prompt-monitoring tools that claim to track AI visibility are reliable, given noise from personalization and account variance, versus page-level structural analysis based directly on the GEO paper's signals. The founder confirmed the startup applies the paper's statistics/sources/citations techniques directly.
Ghulam Mustafa
About the author
Ghulam Mustafa
Founder

Ghulam Mustafa is the founder of AI Rankings and CEO of a digital marketing agency based in Abu Dhabi, UAE. His career sits at the intersection of full-stack development and search — building on Flask, Django, WordPress, and JavaScript while running SEO, AEO, and GEO campaigns for clients across the region. AI Rankings grew out of that work: a platform for tracking how brands actually show up in AI-generated answers, built on the principle that every number it reports has to be real and verifiable, never estimated or simulated. He writes about AI search visibility, technical SEO, and the shift from ranking on Google to being cited by AI.

View profile →

See where you actually stand right now.

Free, live check. Real evidence, not an estimate.

Run your free AI visibility check Talk to us instead