Blog

What Generative Engine Optimization Research Shows: Only Two Tactics Survived a 45-Study Audit

The most comprehensive review of generative engine optimization research to date, a July 2026 audit of 45 studies published between November 2023 and July 2026, found that only two tactics show a reproducible effect on whether AI answers cite a page: topical relevance to the query and where a source sits within the retrieved context. Every other widely repeated GEO claim, including the claim that optimization lifts visibility by 40 percent, either lacked supporting evidence or was formally rejected by the review.

The paper is called "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023 to 2026)", written by Olivier Martinez and posted to arXiv on July 15, 2026. It has not gone through peer review yet, and it is not the kind of study that gets forwarded around a marketing Slack channel. But it is the first attempt to grade the entire body of generative engine optimization research against a consistent evidence standard, and the grades it hands out contradict a lot of what gets sold as GEO best practice.

Key takeaways

  • A 45-study academic audit of generative engine optimization research found only two tactics with reproducible, causal evidence: topical relevance and context position.
  • The often-cited "GEO increases visibility by 40 percent" statistic is a rejected claim. It is a relative maximum on one metric (Position-Adjusted Word Count) under one experimental configuration, not a real-world traffic or citation lift.
  • Rewriting body copy to sound more citable can backfire: a 171,003-document benchmark found body-only rewrites cut a page's presence in the top 10 retrieved results, after reranking, by 16 percent.
  • Structured, extractable content does the opposite of a citation-chasing rewrite: structural optimization alone improved retrieval hit rate by 22 percent in the same benchmark.
  • No study reviewed established a stable link between AI citations and clicks, conversions, or revenue, which is why measurement discipline matters more than any single tactic right now.

See how your brand shows up in AI answers, free.

How the generative engine optimization research was conducted

Martinez's survey does not treat GEO as a single skill. It breaks the problem into eight pipeline stages: activation (whether an AI search is triggered at all for a given query), crawling and indexing, retrieval, reranking and context allocation, generation and citation, absorption and fidelity, attention and click behavior, and downstream economic outcomes. Most public GEO advice, including a lot of what gets published under "GEO checklist" headlines, only addresses one of those eight stages: generation and citation, the moment an AI model decides whether to reference a source it has already retrieved.

That framing matters because it explains why so much generative engine optimization research disagrees with itself. A tactic that improves citation rate for content already sitting in a model's context window says nothing about whether that content gets retrieved in the first place. The survey's central finding is blunt about this gap: within the reviewed evidence, "already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior."

In plain terms: researchers can reliably move the needle on what happens after a page is already in the running. Nobody has produced durable evidence that a specific GEO technique gets a page into the running more often, across engines, over time.

The claim the survey rejected outright

The most quoted statistic in GEO marketing content is that optimization increases visibility by 40 percent. It traces back to the original 2023 GEO paper by Aggarwal et al., which introduced the term and tested nine content interventions (adding statistics, citing sources, quotation, and similar) against a metric called Position-Adjusted Word Count, or PAWC.

Martinez's survey traces that 40 percent figure to its source and rejects it as a general claim. The number reflects a Position-Adjusted Word Count increase from roughly 19.3 to 27.2 under one specific configuration in a controlled test set, where the source document was already placed in a fixed context window. It measures word-count-weighted prominence within an answer that already cited the page, not the odds that the page gets retrieved and cited to begin with. Treating it as an expected traffic or citation lift, which is how it circulates in sales decks and blog posts, is not what the original data supports.

The two tactics with reproducible evidence

Two findings held up across multiple independent studies and made it into the survey's "high confidence" tier.

Topical relevance. How closely a document's content aligns with the literal question a user asked outweighs most other signals researchers have tested. The survey notes that generative models "strongly favor explicit alignment with the question, often more than human credibility cues such as scientific references or a neutral tone." That is a meaningful correction for teams assuming that E-E-A-T signals carry into AI generation the same way they carry into classic search ranking.

Context position. Where a source sits within the set of documents a model is given to work from changes citation odds more than most on-page rewrites do. The survey states plainly that "moving a source higher in the context has a greater effect than most rewrites," which puts the emphasis back on retrieval and reranking, the stages before generation, rather than on copy tweaks aimed at the model itself.

Both of these are upstream of the writing. They are closer to information architecture and retrieval engineering than to prompt-flavored copywriting, which is a real shift from how GEO gets marketed.

Where rewriting for AI citations can backfire

The survey's most useful practical warning comes from SAGEO Arena, a benchmark built on 171,003 documents and 2,700 queries that isolates which part of a page (structural markup versus body text) drives which pipeline stage. The results argue against the common advice to rewrite body copy into FAQ-style, citation-bait paragraphs.

Optimization scope Retrieval hit rate Average rank change Top-10 presence after reranking Final citation
Structural elements only (schema, headings, metadata) +22% +2.72 positions Not degraded Not degraded
Body text only (citation-style rewrites) -9% -4.54 positions -16% -6%
Both together Gains diminish versus structural-only Mixed Mixed Mixed

Structured, query-dense markup is what gets a document through the retrieval gate in the first place, largely because it increases lexical overlap with the terms a BM25-style retriever is matching against. Body text still matters, but at a later stage: it is what a model draws on once a document has already been retrieved and placed in context. Optimizing the body in isolation, in an attempt to sound more quotable, measurably hurt this benchmark's retrieval and reranking performance before a model ever got the chance to cite the page. Combining both did not add the gains together in a straightforward way either. The interaction between the two diluted the structural benefit.

This lines up with what we see in entity and schema work for clients: the unglamorous technical layer (structured data, clear headings, consistent entity naming) tends to do more for brand entity optimization than any amount of body copy written to sound "AI-friendly."

The missing link: citations still don't predict revenue

The survey's lowest confidence grade goes to the claim that citation scores predict clicks, conversions, or revenue. It found only "one suggestive quasi-experiment and a few industry claims," with no established causal link. That is a hard thing for anyone running a GEO program to hear, because it means the metric most vendors report (share of citations, mention frequency, visibility score) is not yet proven to correlate with anything a finance team cares about.

It does not mean measurement is pointless. It means measurement has to stay honest about what it can and cannot claim. Tracking whether a brand is cited is still the only way to know if it is present in AI answers at all, which is the precondition for everything downstream. Our own framework for tracking brand mentions in AI search treats citation tracking as a leading indicator, not a revenue metric, for this reason.

What this means for GEO priorities

Turning the survey's confidence grades into a working checklist gives a clearer picture of where effort belongs right now.

Confidence grade What the research supports What to do about it
High Query relevance and context position drive citation more than most other signals Prioritize retrieval: structured data, clear query-matching headings, and content that answers the literal question first
High Commercial AI engines differ from each other and change over time Test per engine, not once; re-test after model or ranking changes instead of treating one audit as permanent
Moderate Extractable, well-structured content is used more often, depending on intent and factuality Keep structure and factual precision as baseline hygiene, not a silver bullet
Low A single "white-hat" GEO change durably improves discoverability across engines Do not promise clients or leadership a lasting lift from any one tactic
Lowest Citation volume predicts clicks or revenue Report citation tracking as a leading indicator, and pair it with actual referral and conversion data before claiming ROI
Rejected GEO delivers a fixed percentage visibility lift (such as 40 percent) Drop the stat. Ask any agency or tool that cites it to show the underlying methodology

This is a useful gut check when evaluating outside help too. If an agency's GEO services pitch leans on a fixed visibility percentage or promises citation gains without touching retrieval and structured data, that pitch is standing on the same rejected evidence this survey walked back. A credible program starts with retrieval-stage fundamentals, tests per engine, and reports citation data with the caveats the research supports.

If you want a read on where your own brand stands against these two proven levers, an AI Visibility Strategy Call walks through current citation and retrieval performance across the major assistants before recommending any specific fix.

Frequently asked questions

Is generative engine optimization real, or is it hype?

Both, depending on the claim. The underlying mechanism (AI models retrieve and cite some sources more than others, and specific factors influence that) is real and demonstrated in controlled studies. What is hype is the idea that any single rewrite technique produces a guaranteed, durable visibility lift across engines. The evidence supports narrower, upstream levers: relevance to the query and position in the retrieved set, not broad promises.

What is the most reliable GEO tactic according to research?

Topical relevance to the user's query and a document's position within the retrieved context set are the two tactics with the strongest reproducible evidence, according to the 45-study survey. Both sit upstream of on-page copywriting, closer to retrieval and information architecture than to rewriting sentences to sound more quotable.

Does GEO increase visibility by 40 percent?

No. That figure comes from the original 2023 GEO paper's Position-Adjusted Word Count metric moving from about 19.3 to 27.2 under one specific test configuration where the source was already placed in the model's context. The 2026 survey formally rejects it as a general visibility or traffic claim, since it measures word-count-weighted prominence within an existing citation, not the odds of getting cited in the first place.

Should brands stop doing GEO after this survey?

No. The survey argues for a more disciplined version of GEO, not for abandoning it. Structured data, query-matching content, and per-engine testing all have supporting evidence. What should stop is treating single rewrite tactics as guaranteed wins, and reporting citation counts as if they are proven proxies for revenue.

The takeaway

Generative engine optimization research is finally catching up to the marketing built on top of it, and the gap between the two is wider than most practitioners assumed. The tactics with real evidence behind them (query relevance, context position, structured retrieval signals) are less exciting than a 40 percent visibility promise, but they are the ones a brand can defend when someone asks for proof. Build a program around those, measure citation as a leading indicator rather than a revenue metric, and treat every other claim as unproven until a study like this one says otherwise.

See how AI describes your brand

Get a free AI Visibility Report for your category and competitors.