Adding statistics, quotations, and citations to credible sources are the content changes that most reliably increase how often an AI answer pulls from your page, lifting visibility by up to about 40% in the foundational GEO experiment, while keyword stuffing made a page roughly 10% less visible. That finding comes from "GEO: Generative Engine Optimization", a 2024 study by Princeton University researchers (Aggarwal and colleagues, arXiv 2311.09735) that tested nine content edits against a simulated answer engine across 10,000 queries.
It is the closest thing generative engine optimization has to a controlled experiment, so it is worth reading directly. The numbers below are relative gains inside a benchmark, not live-engine guarantees, and we flag the caveats plainly. But the ranking of what works, evidence over repetition, has held up.
What did the Princeton GEO study test?
The researchers built GEO-bench, a benchmark of 10,000 queries across 9 datasets and 7 domains, and simulated a two-stage generative engine: a search step retrieved the top five sources for a query, then GPT-3.5-turbo synthesised an answer that cited them. They then applied nine content edits to a source, one at a time, and measured whether the edited page showed up more prominently in the synthesised answer.
Two custom metrics carried the measurement, and both matter because a raw citation count misses how much of your content the answer actually used:
- Position-Adjusted Word Count — how much of your source's wording appears in the answer, weighted by how prominently it is cited. It rewards being quoted at length and early, not just linked.
- Subjective Impression — a model-rated score of how relevant, influential, and prominent your source feels within the answer, closer to how a reader would judge whose voice dominated.
The setup is a laboratory, not the live web. That is a strength for isolating cause and a limit for reading the exact magnitudes, and we return to it below.
Which changes lifted AI visibility the most?
The evidence-adding edits won; the repetition-based ones lost. Here are the load-bearing results, all relative to an unedited baseline and all from the Princeton paper:
| Content edit | Reported effect on visibility | Read as |
|---|---|---|
| Add statistics (specific numbers) | ~+41% Position-Adjusted Word Count; ~+37% Subjective Impression | Strongest single edit |
| Cite sources (link authoritative references) | Up to ~+115% for a source ranked 5th | Biggest equalizer for lower-ranked pages |
| Add quotations (from credible sources) | ~+28% Subjective Impression | Strong, especially on perceived authority |
| Fluency optimization (clearer writing) + statistics | Beat every single method by ~5.5%+ | Combine, don't pick one |
| Authoritative / easy-to-understand language | Positive, smaller | Supporting edits |
| Keyword stuffing | ~10% worse than baseline | Actively harmful |
The one-line summary: making your claims specific, quoted, and sourced is what moved the needle; padding the page with the target phrase moved it backwards. Across methods and domains, targeted edits produced roughly 22–41% improvements. The step-by-step way to apply them is in the companion how-to, how to add the elements that lift AI citations.
Why do these specific changes work?
Because a generative engine builds an answer by lifting and recombining passages, and it favours passages that are self-contained, verifiable, and easy to attribute. A specific statistic ("cut onboarding to under five minutes") is a clean, liftable unit; a vague claim ("streamlines onboarding") is not. A citation to a credible source and a named quotation both add corroboration, the signal engines lean on to decide a passage is safe to repeat.
That also explains why keyword stuffing backfired. The engine is not counting how often you say the phrase; it is scoring whether your passage is the best evidence for the sub-question. Repetition adds no evidence and reads as low-quality, so it can push you down. This lines up with the broader finding that AI search can't be reliably gamed by manipulation tactics, and with our own guidance to grow AI visibility without spam.
The changes that make an AI answer pull from your page are the same ones that make a human trust it: a specific number, a named source, a clean quotable line. Repetition is not one of them.
How much should you trust these numbers?
Enough to act on the ranking of methods, not enough to quote the exact percentages as if they came from ChatGPT today. The honest limits, several of which the authors and later critics flag:
- Simulated, not live. The engine was GPT-3.5-turbo synthesising Google's top-5 in 2023. Today's engines use newer models, fan a query out into many sub-queries, and pull from far more than five sources.
- A zero-sum, five-source arena. With only five competing sources per query, a gain for one is a loss for another, which amplifies relative percentages. Real niches with hundreds of competing pages will show smaller effects.
- Model-generated edits. The optimizations were written by an LLM, not a human editor, so quality varied.
- Domain-dependent. Gains differed by topic; the paper explicitly calls for domain-specific tuning rather than one universal recipe.
What survives all of that is the mechanism and the pecking order: evidence-dense, well-attributed, cleanly written passages get lifted more than thin or padded ones. Treat the study as a direction and a method, and treat your own measured citation coverage on each engine as the number that actually counts.
How does this fit what we already know?
It supplies the experimental backbone under advice the field mostly reached by observation. Our own six factors behind AI recommendations and how to get cited by AI both lead with specificity, corroboration, and clean extraction; the Princeton experiment is why those work. It also reinforces that GEO is SEO's next layer, not a separate trick: the winning edits are editorial quality, not schema hacks. Note the study tested content edits, not structured data, which is a labelling layer that helps machines parse what things are, a different and complementary lever, not one this experiment measured.
If you want to know whether these edits are actually moving your visibility, that is the loop Buffy Intel closes: it snapshots whether AI engines cite and recommend your brand over time, so you can apply the statistics, quotes, and sources this study validated and then watch your earned citation share respond, engine by engine, instead of trusting a benchmark percentage to hold for you.