Field note

How to tell if a GEO study or stat is trustworthy

A step-by-step way to pressure-test any GEO or AI-visibility claim before you act on it: separate which stage was measured (retrieval, prominence, citation, or traffic), check whether the test was live or simulated, weigh sample size, engines, and dates, and demand corroboration. Built from the questions a July 2026 critical survey of 45 studies says most claims fail.

Buffy Editorial2026-07-27 · 4 min read

To judge a GEO or AI-visibility claim, ask which stage of the pipeline it measured, whether the test was live or simulated, how big and how recent it was, and whether anything independent corroborates it, then act only on findings that survive all four. A July 2026 critical survey of 45 studies (Olivier Martinez, arXiv 2607.14035) found that most headline GEO numbers fail at least one of these checks, usually by measuring a late-stage lab metric and letting readers hear it as a traffic promise. This is the checklist that catches that, whether the claim comes from a vendor, a blog, or a peer-reviewed paper.

Step 1: Which stage of the pipeline did it measure?

Separate the four stages before you read the number, because a gain in one says little about the others:

Stage The question it answers Easy to move?
Retrieval Were you pulled into the candidate set at all? Hard, this is the real bottleneck
Prominence How much of your wording did the answer use? Easy, especially in a lab
Citation Were you named or linked? Medium
Traffic Did a human actually click through? Hardest to earn, rarely reported

Most "+X% AI visibility" claims measure prominence (for example, Position-Adjusted Word Count) and invite you to hear traffic. The survey's critique of the famous 40% figure turns entirely on this gap: the gain was a prominence metric, not a retrieval or click gain. If a claim will not name its stage, stop reading it as a business result.

Step 2: Was it a live engine or a simulated one?

Find out what the study actually tested against. A simulated engine synthesises an answer from a small, fixed set of documents the researcher supplies. That is excellent for isolating cause, but it removes the hardest part of real visibility: getting retrieved from millions of pages in the first place. A live test queries ChatGPT, Perplexity, Google AI Mode, or AI Overviews as they actually behave, with query fan-out, reranking, and real competition.

Ask: how many competing sources were in play? A five-source, zero-sum testbed inflates relative percentages, because one source's gain is mechanically another's loss. Real niches with hundreds of competing pages show far smaller effects. Treat simulated numbers as direction; trust live, measured numbers for magnitude.

Step 3: How big, how many engines, and how recent?

Weigh the sample and the shelf life, in this order:

  1. Sample size. A handful of prompts or one brand per vertical is a signal, not a law. Ask how many queries, how many pages, how many brands.
  2. Engine coverage. A finding on one engine rarely transfers; the survey found generic heuristics transfer poorly, and citation overlap between engines is low. "Works on AI search" usually means "worked on one engine, once."
  3. Dates, not just publication date. Check when the test happened and which model versions it used. AI engines change monthly, so a six-month-old citation study may describe a world that no longer exists, this is the freshness cliff applied to research itself.

Step 4: Does anything independent corroborate it?

Demand a second source before you act. A lone result, especially a single vendor reporting on its own product, is a hypothesis. Corroboration from an independent method or dataset is what moves it toward fact. Two useful tells:

  • Direction of incentive. Does the party publishing the number sell the thing the number promotes? That does not make it false, but it raises the bar for corroboration.
  • Reproducibility. Has anyone repeated the test and seen the same effect? The survey explicitly calls for repeated measurements, paraphrase controls, and human validation, if a study did none of these, hold its magnitude loosely.

A GEO claim earns your action when it names the stage it measured, was tested on live engines at real scale, is recent, and is corroborated independently. Anything short of that is a direction to test, not a tactic to deploy.

Step 5: Test it against your own measured visibility

End every appraisal by checking the claim on your own site, because even a well-built finding may not hold for your niche. Apply the change to a subset of pages, then track your citation coverage per engine before and after, watching retrieval and citation, not just prominence, so you catch the backfire case where a citation-friendly rewrite makes a page harder to retrieve. The broader method for that measurement is in how to measure AI visibility and how to audit your site for AI visibility.

This is the discipline Buffy Intel is built around: it snapshots whether AI engines cite and recommend your brand across engines over time, so the GEO claims you keep are the ones your own data confirms, and the ones that failed this checklist quietly get dropped before they cost you.

Frequently asked

What is the single most important question to ask about a GEO stat?

Which stage did it measure? AI visibility is a pipeline, retrieval (were you pulled into the candidate set), prominence (how much of you the answer used), citation (were you linked), and traffic (did anyone click). A figure like '+40% visibility' is meaningless until you know which of those four it describes. Most inflated claims quietly measure prominence, the easiest to move in a lab, and let readers assume it means traffic, the hardest to earn. If a source will not tell you the stage, treat the number as marketing.

Are simulated GEO experiments worthless?

No, but they answer a narrower question than they appear to. A simulated engine with a fixed set of pre-supplied sources is good for isolating cause, does edit X change how prominently source Y appears? It cannot tell you whether X helps you get retrieved on a live engine with millions of competing pages, query fan-out, and reranking. Use simulated results for direction and mechanism; use your own measured citation data on live engines for magnitude and whether it holds for you.

How fresh does a GEO study need to be?

Fresher than most fields, because the engines change monthly. A citation-pattern study from more than a few months ago may describe models and retrieval behavior that no longer exist. Always check the test dates (not just the publication date), which engines and model versions were tested, and whether anyone has reproduced it since. Pair any single study with your own current measurement rather than assuming last quarter's finding still holds.