To judge a GEO or AI-visibility claim, ask which stage of the pipeline it measured, whether the test was live or simulated, how big and how recent it was, and whether anything independent corroborates it, then act only on findings that survive all four. A July 2026 critical survey of 45 studies (Olivier Martinez, arXiv 2607.14035) found that most headline GEO numbers fail at least one of these checks, usually by measuring a late-stage lab metric and letting readers hear it as a traffic promise. This is the checklist that catches that, whether the claim comes from a vendor, a blog, or a peer-reviewed paper.
Step 1: Which stage of the pipeline did it measure?
Separate the four stages before you read the number, because a gain in one says little about the others:
| Stage | The question it answers | Easy to move? |
|---|---|---|
| Retrieval | Were you pulled into the candidate set at all? | Hard, this is the real bottleneck |
| Prominence | How much of your wording did the answer use? | Easy, especially in a lab |
| Citation | Were you named or linked? | Medium |
| Traffic | Did a human actually click through? | Hardest to earn, rarely reported |
Most "+X% AI visibility" claims measure prominence (for example, Position-Adjusted Word Count) and invite you to hear traffic. The survey's critique of the famous 40% figure turns entirely on this gap: the gain was a prominence metric, not a retrieval or click gain. If a claim will not name its stage, stop reading it as a business result.
Step 2: Was it a live engine or a simulated one?
Find out what the study actually tested against. A simulated engine synthesises an answer from a small, fixed set of documents the researcher supplies. That is excellent for isolating cause, but it removes the hardest part of real visibility: getting retrieved from millions of pages in the first place. A live test queries ChatGPT, Perplexity, Google AI Mode, or AI Overviews as they actually behave, with query fan-out, reranking, and real competition.
Ask: how many competing sources were in play? A five-source, zero-sum testbed inflates relative percentages, because one source's gain is mechanically another's loss. Real niches with hundreds of competing pages show far smaller effects. Treat simulated numbers as direction; trust live, measured numbers for magnitude.
Step 3: How big, how many engines, and how recent?
Weigh the sample and the shelf life, in this order:
- Sample size. A handful of prompts or one brand per vertical is a signal, not a law. Ask how many queries, how many pages, how many brands.
- Engine coverage. A finding on one engine rarely transfers; the survey found generic heuristics transfer poorly, and citation overlap between engines is low. "Works on AI search" usually means "worked on one engine, once."
- Dates, not just publication date. Check when the test happened and which model versions it used. AI engines change monthly, so a six-month-old citation study may describe a world that no longer exists, this is the freshness cliff applied to research itself.
Step 4: Does anything independent corroborate it?
Demand a second source before you act. A lone result, especially a single vendor reporting on its own product, is a hypothesis. Corroboration from an independent method or dataset is what moves it toward fact. Two useful tells:
- Direction of incentive. Does the party publishing the number sell the thing the number promotes? That does not make it false, but it raises the bar for corroboration.
- Reproducibility. Has anyone repeated the test and seen the same effect? The survey explicitly calls for repeated measurements, paraphrase controls, and human validation, if a study did none of these, hold its magnitude loosely.
A GEO claim earns your action when it names the stage it measured, was tested on live engines at real scale, is recent, and is corroborated independently. Anything short of that is a direction to test, not a tactic to deploy.
Step 5: Test it against your own measured visibility
End every appraisal by checking the claim on your own site, because even a well-built finding may not hold for your niche. Apply the change to a subset of pages, then track your citation coverage per engine before and after, watching retrieval and citation, not just prominence, so you catch the backfire case where a citation-friendly rewrite makes a page harder to retrieve. The broader method for that measurement is in how to measure AI visibility and how to audit your site for AI visibility.
This is the discipline Buffy Intel is built around: it snapshots whether AI engines cite and recommend your brand across engines over time, so the GEO claims you keep are the ones your own data confirms, and the ones that failed this checklist quietly get dropped before they cost you.