Field note

Does GEO actually work? What a critical survey of 45 studies found

A July 2026 critical survey of 45 GEO studies (2023-2026) concludes that no reviewed technique shows a stable, cross-platform effect on whether AI engines retrieve or send traffic to a page, and that the famous 40% visibility uplift is a post-retrieval prominence gain measured in a testbed where the source was already supplied. Here is what the evidence actually supports, and where content-tweak GEO stops.

Buffy Editorial2026-07-27 · 5 min read

The evidence that specific GEO content tweaks reliably increase whether AI engines retrieve or send traffic to your page is weak; the effects that replicate are narrower than the marketing suggests. That is the conclusion of a July 2026 critical survey of 45 studies (2023-2026) by Olivier Martinez, "Optimizing Visibility in Generative Engines" (arXiv 2607.14035, dated 15 July 2026). Its finding, in the authors' words: "already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." It is a single-author pre-print, so read it as a rigorous synthesis to pressure-test, not settled law. But its core distinction is one every generative engine optimization claim should be held to.

What did the survey actually find?

That GEO is not one ranking task but a pipeline of stages, and evidence for a technique at one stage rarely carries to the others. The survey models the path as: search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, factual absorption, and finally user behavior. A content edit can help late stages (how much you are quoted once selected) while doing nothing, or worse, at the early stage that decides whether you are selected at all.

Within that frame, three results are worth pinning down, all attributed to the survey and its reviewed corpus:

Claim What the evidence shows How to read it
Content edits lift visibility Real, but only after retrieval, in controlled testbeds Prominence effect, not a retrieval or traffic effect
A single technique reliably wins across engines Not found in 45 studies Generic heuristics transfer poorly between engines
More citation-optimized = more cited overall Sometimes the reverse Rewrites can cut retrieval, negating downstream gains

The one-line summary: the most reproducible levers were topical relevance and context position, not clever formatting; and competition erodes individual gains as everyone applies the same edits.

Where does the "40% uplift" number come from?

From a single metric in one controlled experiment, not from live traffic. The survey traces the widely-quoted 40% to the foundational Princeton GEO study, covered in our own explainer on which content changes lift AI citations. There, the Position-Adjusted Word Count metric rose from 19.3 to 27.2 (about a 41% relative gain) under a "Quotation Addition" strategy.

The catch the survey stresses: that gain occurred "conditional on a source already being present in a fixed context", a testbed where five documents had already been supplied to the generator. In the survey's words, the foundational gains "are valid within its experimental setting" but "establish neither organic discoverability nor durable traffic effects." So the number describes how much more of your wording an answer reuses once you are in the candidate set. It says nothing about your odds of getting into that set, and nothing about clicks.

Can GEO edits make a page harder to find?

Yes, and this is the survey's most useful correction. It reports an end-to-end test (the SAGEO Arena setup) in which optimizing only a page's body text for citability:

  • cut top-10 presence after reranking by about 16%
  • reduced average top-20 presence by roughly 9%
  • lowered final citation by about 6%

A rewrite that makes your passage more quotable once retrieved can make the whole page a worse match at retrieval, so it is pulled into the answer less often. Optimizing the prominence stage can quietly sabotage the retrievability stage.

This is why "we added statistics and quotations, so we will be cited more" does not always hold on a live engine. The retrieval step and the prominence step reward partly different things, and a change tuned for one can cost you the other. The step-by-step way to check a claim like this before you act on it is in the companion how-to, how to tell if a GEO study or stat is trustworthy.

Does this contradict the Princeton GEO study, or Buffy's own advice?

No, and holding both without contradiction is the point. The Princeton experiments measured a prominence metric in a simulated engine and were careful to call the numbers controlled-environment relatives; our write-up already flagged the zero-sum, five-source arena as the reason the percentages inflate. The Martinez survey does not overturn that, it formalizes the caveat and quantifies the missing stages: prominence is not retrieval, and retrieval is not traffic.

It also lines up with Google's own position that, on its surfaces, GEO is "still SEO", no special files or tricks, just genuinely useful, crawlable, corroborated content. Two independent sources, an academic survey and a search engine, arriving at the same unglamorous answer is corroboration, the signal worth trusting. The takeaway is not "GEO is fake." It is that durable levers beat technique-chasing, and the only figure that counts for you is your own measured citation coverage per engine, which is exactly what a content edit's effect should be validated against rather than a benchmark percentage.

What should you actually do about it?

Stop buying single techniques as guarantees, and start treating visibility as a measured, multi-stage outcome:

  1. Compete on relevance and evidence, not formatting hacks. The reproducible levers were topical relevance and clear, corroborated substance, the same things that survive freshness decay.
  2. Test edits against retrieval, not just prominence. Before rolling a rewrite site-wide, confirm it did not make the page harder to retrieve, the survey shows that failure mode is real.
  3. Measure per engine, over time. Citation behavior differs by engine and shifts weekly; a one-off benchmark cannot tell you if you are winning.
  4. Distrust any "+X% visibility" claim that will not say which stage it measured (retrieval, prominence, citation, or traffic) and on which live engines.

That measurement loop, snapshotting whether AI engines cite and recommend your brand across engines and over time, is what Buffy Intel is built to close. The survey's lesson is Buffy's thesis stated academically: you cannot optimize what you refuse to measure, and the techniques worth keeping are the ones your own citation data confirms, not the ones a lab percentage promised.

Frequently asked

Does GEO actually work, or is it hype?

Both, in different places. According to a July 2026 critical survey of 45 studies (Olivier Martinez, arXiv 2607.14035), content edits can causally change how a source is cited or used once it has already been retrieved into an answer, that part replicates. What does not replicate is any single technique reliably improving whether a page gets retrieved in the first place, or whether it earns real traffic, across engines and over time. So GEO 'works' at the prominence stage and is unproven at the retrieval-and-traffic stage. The durable levers are unglamorous: be genuinely relevant, crawlable, corroborated, and fresh, then measure where you are actually cited.

Is the '40% increase in AI visibility' figure real?

It is real inside its experiment and widely over-read outside it. The survey traces the 40% to the foundational Princeton GEO paper, where the Position-Adjusted Word Count metric rose from 19.3 to 27.2 (about a 41% relative gain) under a 'Quotation Addition' strategy. That was measured in a controlled testbed where five documents had already been supplied to the generator. It describes how much more of your wording an answer uses once you are in it, not a 40% gain in being retrieved, cited, or clicked. Treat it as a prominence result, not a traffic promise.

Can optimizing a page for AI citations backfire?

Yes, the survey documents exactly that. In one end-to-end test, rewriting only a page's body text to be more citation-friendly cut its top-10 presence after reranking by about 16%, reduced average top-20 presence by roughly 9%, and lowered its final citation by about 6%. The mechanism: a rewrite tuned to be quotable once retrieved can make the page a worse match at the retrieval stage, so it is selected less often to begin with. Optimizing one stage of the pipeline can quietly harm another.