Field note

How to test whether a change actually moved your AI citations

Adding schema, refreshing a page, or rewriting an intro either helped your AI citations or it didn't, and a single before-and-after can't tell you which. Here's how to test an on-page change like a controlled experiment: isolate one variable, baseline it, hold out a control set, wait out the crawl-to-cite lag, and read the delta honestly.

Buffy Editorial2026-08-06 · 4 min read

You can't know whether adding schema, refreshing a page, or rewriting an intro moved your AI citations from a single before-and-after, because citations drift on their own. To test it honestly, treat the change like a controlled experiment: isolate one variable, baseline it, hold out a comparable control set, wait out the crawl-to-cite lag, then read the delta against the control, not against zero. This is the single-site version of the method behind large studies like the controlled schema experiment, and it's how you avoid mistaking correlation for causation.

This is the hands-on companion to that reference; read it for the finding, this for the method. Below is the test in six steps.

Step 1: Isolate one change and write the hypothesis

You can only attribute an effect to a change you can name. Before touching anything, write one sentence: "adding Product schema to these 40 product pages will increase their citations on ChatGPT within six weeks." One variable, one page-set, one engine, one window.

The discipline is to change one thing. If you add schema and rewrite the copy and add FAQs in the same week, any movement is unattributable, you'll never know which edit did it. Stagger changes, or accept that you're testing the bundle, not the schema.

Step 2: Baseline before you touch anything

A test needs a starting line. For the target pages and prompts, record where you stand now: which prompts you're cited in, on which engines, and how often, ideally as an average over a couple of weeks so a single volatile reading doesn't become your baseline. This is ordinary AI-visibility measurement, captured deliberately before the change rather than reconstructed after.

Capture the raw detail, not just a score: the specific prompts, the engine, and whether you were cited or merely mentioned. You'll need that granularity in Step 5 to tell a real effect from a reshuffle.

Step 3: Hold out a control set

This is the step that separates a real test from a guess. Pick a set of comparable pages, similar topic, similar current citation level, and don't change them. They're your control. Both groups live through the same engine updates, the same season, the same competitor moves, so whatever happens to the control is the background drift you must subtract.

  • Treated group: the pages you change.
  • Control group: similar pages you deliberately leave alone.
  • The signal: how much the treated group moves beyond the control, not how much it moves in absolute terms.

Without a control, a rise that's really just an ecosystem-wide good week gets miscredited to your change, the exact trap a naive before-and-after falls into.

The question is never "did citations go up after I made the change?" It's "did the pages I changed move more than the near-identical pages I didn't?" The control is what turns a story into evidence.

Step 4: Wait out the crawl-to-cite lag

Give the change time to propagate. A page must be recrawled, reprocessed, and then selected at answer time, so effects show up in weeks, not days. Reading the result on day three mostly measures answer volatility, not your edit.

Watch a smoothed trend over several weeks rather than any single snapshot. Heavy crawler activity with no citation movement yet is normal and expected, crawl volume tends to rise before citations do, so don't call the test dead in week two. Set the judging date in advance so you're not tempted to stop the moment the line ticks the way you hoped.

Step 5: Read the delta honestly, and rule out confounders

Now compare. Did the treated group's citations rise relative to the control over the window? Before you conclude "yes," rule out the usual confounders:

  • Other changes: did anything else ship to the treated pages, a template update, an internal-linking change, a freshness bump, in the same window?
  • Composition: are you counting citations, or has a mention simply been reclassified? Keep the two apart.
  • Sample size: on a handful of pages, a two-citation swing is noise. Read the pattern across many prompts and pages, not one URL.
  • Direction of cause: if the change and the citations both reflect a third factor (you refreshed your best pages), you've found a correlation, not a cause.

A move that survives all four is a credible, site-specific result. One that doesn't is a lead to investigate, not a conclusion to ship.

Step 6: Decide, and size the claim to the evidence

Turn the result into a right-sized decision. If the treated group beat the control convincingly, roll the change out and keep watching. If it didn't, stop spending on that change, this is precisely how the schema study's null result should redirect effort toward the levers that do move citations. Either way, state the claim at the confidence your sample supports: "on our site, for these pages, over six weeks" is honest; "schema works" from forty pages is not.

Testing changes this way, one variable, against a control, over a real window, is how you learn what actually moves your citations instead of copying tactics that were only ever correlations. Sustaining that measurement continuously, across every engine and over time, is exactly what Buffy Intel is built to do. Change one thing; prove it moved you. Questions: [email protected].

Frequently asked

Why can't I just compare citations before and after I made a change?

Because AI citations drift on their own, so a naive before-and-after confuses your change with everything else that moved in the same window: an engine update, a competitor's new page, seasonality, or normal answer volatility. If citations rise after you add schema, you can't tell whether the schema did it or the engine simply re-ranked that week. The fix is a control: track comparable pages you did not change over the same period, so the difference between the two isolates your change from the background drift.

How long should I wait before judging whether a change worked?

Longer than feels natural, because there's a crawl-to-cite lag: a page has to be recrawled, reprocessed, and then chosen at answer time, which can take weeks. Judging in the first few days mostly measures noise. Give it several weeks, read a smoothed trend rather than any single reading, and remember that heavy crawler activity with no citation change yet is a normal leading indicator, not a failed test.

Do I need a big sample like the published studies?

For a confident causal claim, yes, one page proves almost nothing, which is exactly why single-vendor studies use thousands of matched pages. On a single site you usually can't reach that, so treat your test as directional evidence for your own site, not proof of a general rule. Group similar pages, use a holdout set, and read the pattern across many prompts and pages rather than betting on one URL.