Field note

Why your AI 'share of voice' score swings week to week

Ask an AI the same category question 30 times and you get the same brand list only 50-61% of the time; reword it slightly and the overlap collapses below 30%. A 2026 study quantifies how volatile AI answers are, and why a single citation-rate snapshot is the wrong thing to track. Here is the data, attributed, and what to measure instead.

Buffy Editorial2026-06-30 · 6 min read

AI answers are probabilistic, not fixed. Ask an AI engine the same category question repeatedly and the set of brands it names changes; reword the question slightly and the list scrambles further. So a single snapshot of your "share of voice" or citation rate is volatile by design: it moves week to week even when nothing about your brand changed. The fix is not to abandon measurement but to change the method: sample many prompts, repeatedly, across every engine, and anchor on downstream outcomes.

This piece collects the evidence. Every figure attributed and dated, and explains what to measure instead. It pairs with the deep-dive on how ChatGPT picks the sources it cites and defines the underlying concept in the glossary entry for answer volatility.

How unstable is an AI "share of voice" number, really?

Very. A 2026 study by Zaprev (an India- and US-based AI-visibility vendor), summarised in a widely shared LinkedIn post by founder Mayank Shrivastava, asked roughly 12,000 commercial queries across OpenAI and Anthropic models and measured how much the list of recommended brands overlapped between two answers. It scored overlap with Jaccard similarity: the count of shared brands divided by the combined pool, where 1.0 is identical and 0 is no overlap.

The findings, as reported, are stark. Treat the exact numbers as one vendor's self-reported study (single methodology, shared via social post, not peer-reviewed). Directional rather than precise.

What changed between the two answers Brand-list overlap
Nothing: identical question asked 30 times in one day same set only 50-61% of the time
Synonym swap: "best CRM" → "top CRM" ~28% overlap
Added constraint: "CRM" → "CRM for a SaaS startup under 50 people" ~13.5% (Jaccard)
Region / language: English → French, UK → German ~13.8% overlap
Different provider: same question to OpenAI vs Anthropic ~33% consensus

Source: Zaprev / ZapRank study (2026), as reported by the author on LinkedIn. The headline reading: even the best case. The identical question, asked twice, same day. Agrees only about half the time, and any cosmetic rewording scrambles the brand list more than switching to an entirely different model provider does. Wording moves the answer more than the model does.

Why do AI answers change when nothing about your brand did?

Because four independent sources of variation stack on top of each other, and none of them is under your control:

  • Sampling. Language models generate by sampling tokens probabilistically, so the same prompt yields a slightly different answer each time. Different phrasing, and often a different shortlist.
  • Live retrieval. When the engine searches the web, it pulls a fresh mix of pages each time; what is indexed and judged relevant shifts day to day. (How that selection happens is the subject of how ChatGPT picks sources.)
  • Query fan-out. A single question is silently expanded into many sub-queries, and small wording changes route to different sub-queries and different sources. See how query fan-out works.
  • Phrasing sensitivity. As the table shows, adding a constraint or translating the prompt produces a substantially different answer, because it changes which slice of the corpus the model retrieves against.

The phenomenon is corroborated beyond the single study. An academic preprint, "Don't Measure Once: Measuring Visibility in AI Search" (arXiv, 2026), tracked AI answers over ~45 days and found day-to-day overlap of cited sources averaging only about 0.34-0.42 (Jaccard). Roughly 60-65% of cited sources changing between consecutive days. Vendors acknowledge it too: HubSpot shipped a free "AEO Sensor" specifically to track answer-engine volatility, and Profound publishes an ongoing AI-search-volatility tracker. The instability is a property of the medium, not a flaw in any one tool.

Does a smarter model or "more reasoning" fix it?

No. The same study reported that changing the model's reasoning effort from high to low moved the overlap numbers by at most about 5%: a rounding error next to the 40-50-point swings caused by rewording. You cannot buy your way out of volatility with a bigger model, because the variation lives in sampling and retrieval, not in reasoning depth. This is why a "we asked the smartest model once" snapshot is no more trustworthy than any other single observation.

Does this mean AI-visibility tracking is pointless?

No. It means a single snapshot of a tiny fixed prompt set is unreliable, which is a sampling problem with a sampling answer. The same statistics that make one observation noisy make a large, repeated sample stable. The honest method:

  • Sample many prompts, not a handful. Ten hand-picked prompts will whipsaw; a few hundred representative ones average out. This is why choosing the right prompts to track matters more than any single result.
  • Repeat the sampling on a schedule and read the distribution, not one figure. A share of voice of "42% ± 6 over 300 prompts this week" is a real signal; "43% on Tuesday" is noise.
  • Measure across every engine. Cross-provider consensus was only ~33%, so a number from one engine says little about another. Citation coverage and presence have to be read per engine.
  • Watch the trend, not the wobble. A four-week moving line tells you whether you are gaining ground; a day-to-day delta mostly tells you the model sampled differently. This is exactly the discipline behind reading an AI-visibility case study honestly.

Done this way, presence metrics are useful leading indicators. Done as a single spot check, they are theatre.

What should you actually measure?

Anchor on downstream outcomes, because they aggregate over many answers and tie to real behaviour rather than to one volatile generation. The scoreboard has already moved from clicks to citations, but the most stable layer sits one step further down the funnel:

  • AI-referred sessions and conversions: build an AI-referral view in GA4, remembering that much AI traffic arrives unlabelled as Direct.
  • Pipeline from AI-discovered buyers: the method for tying revenue to AI answers is in how to prove AI traffic converts.
  • Smoothed presence trends as the leading indicator that feeds the above, never a single citation-rate reading.

Treat one AI answer like one coin flip: a single result tells you almost nothing, but ten thousand results tell you the odds. Measure the distribution and the downstream outcome, not the snapshot.

The durable takeaway survives whatever the exact 2026 numbers turn out to be: AI answers are noisy by construction, so a stable reading comes from volume, repetition, and outcomes, not from a prettier single number. Confusing a noisy snapshot for a trend is the most common mistake in AI-visibility reporting.

Reading presence as a smoothed distribution across hundreds of prompts and every major engine, and tying it to AI-referred traffic and conversions. Is exactly what Buffy Intel is built to do: it samples repeatedly over time so you track the signal, not the noise.

Frequently asked

Why does my AI share-of-voice number change when nothing about my brand changed?

Because AI answers are probabilistic, not fixed. The model samples a slightly different response each time, retrieval pulls a different mix of fresh pages, and any change in wording reshapes the answer. A 2026 study (Zaprev, reported via LinkedIn) found that asking an identical question 30 times in one day returned the same brand set only 50-61% of the time, so a single snapshot of citation rate or share of voice moves on its own, independent of anything you did.

Does a single citation-rate snapshot mean my AI-visibility tracking is broken?

It means a single snapshot of a tiny fixed prompt set is unreliable, not that measurement is pointless. The fix is method: sample many prompts (not a handful), repeat the sampling on a schedule, average across enough observations to read a distribution rather than one number, do it across every engine, and watch the trend instead of the week-to-week wobble. Volatility is a sampling problem with a sampling answer.

What is the most stable AI-visibility metric to track?

Downstream outcomes are the most stable signal, because they aggregate over many answers and tie to real behaviour: AI-referred sessions, conversions, and pipeline from buyers who arrived via an AI answer. Presence metrics (share of voice, citation rate) are still useful as leading indicators, but read them as smoothed trends across a large prompt set and many engines, never as a single exact figure.