
The GEO Paper: What the Research Behind GEO Actually Found
- Alan Turkmen

- Aug 18
- 7 min read
Quick answer: "The GEO paper" usually refers to GEO: Generative Engine Optimization, a 2024 research paper by a team including Princeton, Georgia Tech, Allen Institute for AI, and IIT Delhi researchers, which introduced the term and tested methods for improving visibility in AI-generated answers. The study found that adding citations, statistics, and quotations to content produced the largest visibility gains in generative engine responses — in some tests, source visibility improved by over 30% relative to unoptimized content. It also built a benchmark called GEO-bench with thousands of queries to measure how content performs across different generative search systems.
Key takeaways - The original GEO paper was posted to arXiv in 2024 and later appeared at KDD, authored by researchers from Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi. - The paper tested nine content optimization methods, including adding citations, quotations, statistics, and simplifying language, then measured which ones changed how often a source was cited or referenced in AI-generated answers. - Citing sources and adding statistics were among the strongest levers the researchers identified; keyword stuffing — a classic SEO tactic — showed little to no benefit in their tests. - The researchers released GEO-bench, a benchmark of queries across multiple topic categories, so other teams could reproduce and extend the findings.
What Is the GEO Paper, Exactly?
The GEO paper is the academic study that coined the term Generative Engine Optimization (GEO) and gave it a testable framework. Before this paper, there was no standard way to measure whether content was actually more likely to be surfaced or cited by an AI system like an LLM-powered answer engine.
The researchers set out to answer a practical question: if traditional SEO is built around ranking in a list of blue links, what does "ranking" even mean when the response is a single generated paragraph with a handful of citations woven in? Their answer was to define visibility as a measurable score — essentially, how much a given source contributes to and is credited within a generative engine's response.
We've covered the basic concept in plainer terms in GEO Meaning: What Generative Engine Optimization Means, but this post focuses specifically on what the underlying research tested and found.
What Methods Did the Researchers Actually Test?
The paper evaluated nine distinct content adjustments, applying each one to existing web content and then measuring the change in visibility within simulated generative engine responses. The methods clustered into a few categories:
Adding citations — attaching sources or references to claims in the content.
Adding quotations — including quoted statements from relevant, credible figures.
Adding statistics — inserting numerical data and figures relevant to the topic.
Simplifying language — rewriting content to be easier to read.
Fluency optimization — improving grammar and sentence flow.
Technical terminology — using more authoritative, domain-specific language.
Unique word usage — increasing vocabulary diversity.
Authoritative tone — rewriting content to sound more confident and expert.
Keyword stuffing — repeating target keywords, a legacy SEO tactic used as a comparison baseline.
The researchers ran these tests across multiple query categories and multiple generative engines to see which changes consistently moved the needle and which ones didn't transfer well across different systems.
Don't skip this: the paper found that keyword stuffing — one of the oldest tricks in traditional SEO — had little to no positive effect on generative engine visibility, and in some cases hurt it. If your current content strategy leans on keyword density, the research suggests that specific habit doesn't carry over to GEO.
Which Tactics Actually Moved the Numbers?
Citing sources and adding statistics produced the most consistent visibility gains across the categories tested in the GEO paper. According to the researchers, adding citations and relevant statistics could improve a source's visibility score by a meaningful double-digit percentage in several test categories, with citation-adding among the top-performing single interventions.
Here's how the tested methods compared directionally, based on the paper's reported results:
Method | Reported effect on visibility |
Adding citations | Strong positive effect |
Adding statistics | Strong positive effect |
Adding quotations | Moderate positive effect |
Technical terminology | Moderate positive effect, varied by topic |
Authoritative tone | Mixed, topic-dependent |
Fluency optimization | Small positive effect |
Unique word usage | Small or inconsistent effect |
Simplifying language | Mixed — helped some categories, not others |
Keyword stuffing | Little to no effect, sometimes negative |
A useful nuance from the paper: effectiveness varied by query category. A method that helped in a "history" or "science" style query didn't always help the same way in a "business" or "how-to" query. The researchers were explicit that content strategy likely needs to adapt per topic rather than applying one universal formula.
What Is GEO-bench and Why Does It Matter?
GEO-bench is the benchmark dataset the paper's authors built to test and compare optimization methods, consisting of thousands of real-world queries pulled from search and question-answering datasets across roughly a dozen topic domains. Before GEO-bench, there wasn't a shared, repeatable way to test how "visible" a piece of content was inside a generative answer — researchers and companies were essentially guessing.
GEO-bench matters for two practical reasons:
It's reproducible. Other researchers and companies can run the same queries against different generative engines and compare results using the same visibility metric, rather than each group inventing its own scoring method.
It's diverse by design. Because it spans multiple topic domains, it captures the finding that no single tactic works everywhere — which is exactly the kind of nuance a business owner should expect when applying this to their own content.
If you're a business owner rather than a researcher, you won't be running GEO-bench yourself. What matters is that the headline numbers you see quoted about GEO — "citations improve visibility by X%" — trace back to a specific, published testing methodology, not a marketing claim. That's worth knowing when someone pitches you on a "proven GEO framework."
How Does the GEO Paper's Findings Compare to Traditional SEO Research?
The GEO paper's findings both overlap with and diverge from decades of traditional SEO wisdom, and the divergence is the more important part for anyone updating their strategy. Traditional SEO has long emphasized keyword placement, backlink volume, and page structure as ranking signals recognized by search engines like Google.
The GEO paper's tests suggest generative engines weigh things differently:
Keyword density, a long-standing SEO lever, showed weak or negative results in the GEO paper's tests.
Source credibility signals — citations, statistics, quoted experts — showed the strongest results, which lines up with how large language models are trained to weigh authoritative, well-supported text.
Readability and fluency still mattered, but less dramatically than citation-based methods.
This doesn't mean traditional SEO is obsolete — search engines like Google still send substantial traffic, and good technical SEO remains part of a sound web presence. It means the tactics for showing up inside an AI-generated answer aren't identical to the tactics for ranking a page. We walk through the practical difference in GEO Explained: How Businesses Show Up in AI Answers.
What Should a Business Owner Actually Do With This Research?
Treat the GEO paper as evidence for a specific set of habits, not a reason to overhaul your entire content strategy overnight. The findings point toward a short, practical checklist you can apply to existing pages before writing anything new:
Audit your top pages for unsupported claims. Find sentences stating a fact or number with nothing backing it up.
Add citations to named, credible sources wherever you make a factual claim — a government agency, an industry association, a published study.
Insert specific statistics relevant to your topic instead of vague statements like "many businesses" or "a growing number of."
Quote a real expert or documented source where it's genuinely relevant, rather than inserting one for the sake of it.
Cut keyword-stuffed phrases that exist only for search engines and read awkwardly to a human.
Check readability — clear, well-structured sentences performed at least as well as dense, jargon-heavy ones in the paper's tests.
One caveat worth stating plainly: the GEO paper tested visibility within simulated generative engine responses using its own benchmark and methodology at the time of publication. Actual AI systems — ChatGPT, Perplexity, Google's AI Overviews — update their models and ranking behavior regularly, and results on any one system today may not match the paper's original findings exactly. Treat the paper as strong directional evidence, not a guaranteed formula, and expect to keep testing.
If you're earlier in the process and want the fundamentals before diving into tactics, What Is GEO? A Step-by-Step Guide to Getting Started is a better starting point than the research paper itself.
Applying Research to Your Own Content Without a Research Team
Most small and mid-size businesses don't have a data science team to replicate GEO-bench testing on their own website, and that's fine — you don't need one to apply the paper's core lessons. What you need is a consistent editorial habit: back up claims with named sources, use real numbers instead of vague ones, and resist the urge to repeat your target keyword for its own sake.
Where this gets harder is at scale — auditing dozens or hundreds of existing pages, tracking which ones show up in AI answers, and rewriting systematically rather than one blog post at a time. That's the kind of work an AI consulting engagement or a technical content audit is built for, and it's a common thread in the AI integration work SFDIFY does for small and mid-size businesses trying to figure out where their existing content stands today.
If you're weighing whether to tackle this yourself or bring in outside help, our breakdown of what an AI consultant actually does covers where that kind of engagement typically starts and what it doesn't replace.
Want a second opinion on whether your website's content is set up to be cited by AI answer engines? Contact SFDIFY to talk through where your site stands today.
Comments