Field note

Which content changes actually increase AI citations? The Princeton GEO experiments

The foundational GEO study, from Princeton researchers, tested nine content edits against a simulated AI answer engine over 10,000 queries and found that adding statistics, quotations, and citations to credible sources lifted a source's visibility by up to about 40%, while keyword stuffing made it roughly 10% worse. Here is what the experiment measured, which changes won, and how much to trust the numbers.

Buffy Editorial2026-07-23 · 5 min read

Adding statistics, quotations, and citations to credible sources are the content changes that most reliably increase how often an AI answer pulls from your page, lifting visibility by up to about 40% in the foundational GEO experiment, while keyword stuffing made a page roughly 10% less visible. That finding comes from "GEO: Generative Engine Optimization", a 2024 study by Princeton University researchers (Aggarwal and colleagues, arXiv 2311.09735) that tested nine content edits against a simulated answer engine across 10,000 queries.

It is the closest thing generative engine optimization has to a controlled experiment, so it is worth reading directly. The numbers below are relative gains inside a benchmark, not live-engine guarantees, and we flag the caveats plainly. But the ranking of what works, evidence over repetition, has held up.

What did the Princeton GEO study test?

The researchers built GEO-bench, a benchmark of 10,000 queries across 9 datasets and 7 domains, and simulated a two-stage generative engine: a search step retrieved the top five sources for a query, then GPT-3.5-turbo synthesised an answer that cited them. They then applied nine content edits to a source, one at a time, and measured whether the edited page showed up more prominently in the synthesised answer.

Two custom metrics carried the measurement, and both matter because a raw citation count misses how much of your content the answer actually used:

  • Position-Adjusted Word Count — how much of your source's wording appears in the answer, weighted by how prominently it is cited. It rewards being quoted at length and early, not just linked.
  • Subjective Impression — a model-rated score of how relevant, influential, and prominent your source feels within the answer, closer to how a reader would judge whose voice dominated.

The setup is a laboratory, not the live web. That is a strength for isolating cause and a limit for reading the exact magnitudes, and we return to it below.

Which changes lifted AI visibility the most?

The evidence-adding edits won; the repetition-based ones lost. Here are the load-bearing results, all relative to an unedited baseline and all from the Princeton paper:

Content edit Reported effect on visibility Read as
Add statistics (specific numbers) ~+41% Position-Adjusted Word Count; ~+37% Subjective Impression Strongest single edit
Cite sources (link authoritative references) Up to ~+115% for a source ranked 5th Biggest equalizer for lower-ranked pages
Add quotations (from credible sources) ~+28% Subjective Impression Strong, especially on perceived authority
Fluency optimization (clearer writing) + statistics Beat every single method by ~5.5%+ Combine, don't pick one
Authoritative / easy-to-understand language Positive, smaller Supporting edits
Keyword stuffing ~10% worse than baseline Actively harmful

The one-line summary: making your claims specific, quoted, and sourced is what moved the needle; padding the page with the target phrase moved it backwards. Across methods and domains, targeted edits produced roughly 22–41% improvements. The step-by-step way to apply them is in the companion how-to, how to add the elements that lift AI citations.

Why do these specific changes work?

Because a generative engine builds an answer by lifting and recombining passages, and it favours passages that are self-contained, verifiable, and easy to attribute. A specific statistic ("cut onboarding to under five minutes") is a clean, liftable unit; a vague claim ("streamlines onboarding") is not. A citation to a credible source and a named quotation both add corroboration, the signal engines lean on to decide a passage is safe to repeat.

That also explains why keyword stuffing backfired. The engine is not counting how often you say the phrase; it is scoring whether your passage is the best evidence for the sub-question. Repetition adds no evidence and reads as low-quality, so it can push you down. This lines up with the broader finding that AI search can't be reliably gamed by manipulation tactics, and with our own guidance to grow AI visibility without spam.

The changes that make an AI answer pull from your page are the same ones that make a human trust it: a specific number, a named source, a clean quotable line. Repetition is not one of them.

How much should you trust these numbers?

Enough to act on the ranking of methods, not enough to quote the exact percentages as if they came from ChatGPT today. The honest limits, several of which the authors and later critics flag:

  • Simulated, not live. The engine was GPT-3.5-turbo synthesising Google's top-5 in 2023. Today's engines use newer models, fan a query out into many sub-queries, and pull from far more than five sources.
  • A zero-sum, five-source arena. With only five competing sources per query, a gain for one is a loss for another, which amplifies relative percentages. Real niches with hundreds of competing pages will show smaller effects.
  • Model-generated edits. The optimizations were written by an LLM, not a human editor, so quality varied.
  • Domain-dependent. Gains differed by topic; the paper explicitly calls for domain-specific tuning rather than one universal recipe.

What survives all of that is the mechanism and the pecking order: evidence-dense, well-attributed, cleanly written passages get lifted more than thin or padded ones. Treat the study as a direction and a method, and treat your own measured citation coverage on each engine as the number that actually counts.

How does this fit what we already know?

It supplies the experimental backbone under advice the field mostly reached by observation. Our own six factors behind AI recommendations and how to get cited by AI both lead with specificity, corroboration, and clean extraction; the Princeton experiment is why those work. It also reinforces that GEO is SEO's next layer, not a separate trick: the winning edits are editorial quality, not schema hacks. Note the study tested content edits, not structured data, which is a labelling layer that helps machines parse what things are, a different and complementary lever, not one this experiment measured.

If you want to know whether these edits are actually moving your visibility, that is the loop Buffy Intel closes: it snapshots whether AI engines cite and recommend your brand over time, so you can apply the statistics, quotes, and sources this study validated and then watch your earned citation share respond, engine by engine, instead of trusting a benchmark percentage to hold for you.

Frequently asked

What is the Princeton GEO study?

It is the 2024 paper 'GEO: Generative Engine Optimization' by Aggarwal and colleagues at Princeton University, presented at the ACM SIGKDD conference (preprint arXiv 2311.09735). It coined the term Generative Engine Optimization and carried out the first large controlled experiment on what content edits make a source more visible inside an AI-generated answer, testing nine methods over a 10,000-query benchmark called GEO-bench. It is the most-cited experimental foundation for GEO, which is why it is worth reading directly rather than through second-hand summaries.

Which content changes increased AI visibility the most?

Adding statistics was the single strongest edit, lifting the paper's Position-Adjusted Word Count metric by about 41% and its Subjective Impression score by about 37%. Adding quotations from credible sources lifted Subjective Impression by roughly 28%, and adding citations to authoritative sources was the biggest equalizer, a source sitting 5th gained about 115% relative visibility. Combining clearer writing with added statistics beat every single method on its own. Keyword stuffing was the notable loser, performing about 10% worse than the unedited baseline.

Do these percentages transfer to live engines like ChatGPT?

Directionally, yes; precisely, no. The study measured a simulated engine (Google's top-5 sources synthesised by GPT-3.5-turbo in 2023) with only five competing sources per query, a zero-sum setup that amplifies relative gains, and the edits were model-generated rather than hand-written. So treat the exact figures as controlled-environment relatives, not promises for a competitive niche on today's models. What is durable is the mechanism: engines lift clean, verifiable, well-attributed passages, and the ranking of methods has held up well against later practitioner data.