Measuring AI VisibilityPart 3 of 9

The AI-visibility reporting stack: 5 metrics that matter

Vanity numbers don't survive a budget review. These are the five metrics that actually tell you whether AI is recommending your brand, and whether it's driving revenue: share of voice, citation rate, sentiment, AI-referral revenue, and a dark-traffic index.

Buffy Editorial2026-06-09 · 4 min read

The fastest way to lose a budget for AI visibility is to report a number nobody trusts. "We're mentioned more!" invites the obvious question, so what?, and a single presence metric can't answer it. To survive a planning review, AI-visibility reporting needs to connect what the engines say to what it drives.

Here's the stack we'd defend: five metrics, two layers. The first three are answer-level: measured inside the AI answers themselves. The last two are outcome-level: measured in your own analytics.

Layer one: what the engines say

These can't be pulled from a rank tracker or GA4. You have to ask the engines the questions your customers ask, across many prompts and every engine, and log what comes back, because answers vary run-to-run, and they disagree with each other.

1. Share of voice

How often you appear, relative to competitors, across a representative set of buyer prompts. This is the headline presence number, but read it as share, not raw count: appearing in 40% of answers means little until you know a competitor is in 80%. Track it per engine, because ChatGPT, Gemini, Claude, and Google's AI surfaces each draw a different shortlist.

2. Citation rate

When you are mentioned, is your own site the cited source, or is the engine learning about you through a retailer, a review site, or a competitor's comparison page? High presence with low citation rate means the narrative about you is being written by others. This is the most actionable metric in the stack, because it points straight at content you can fix.

3. Sentiment

How you're described, not just whether you appear. "Premium and well-reviewed" and "a cheaper alternative" are both mentions; only one helps you. Sentiment turns presence into positioning, and it's where a brand most often discovers the gap between how it sees itself and how the models summarise it.

Answer-level metric Question it answers The trap if you ignore it
Share of voice Do we appear, vs competitors? Counting raw mentions with no benchmark
Citation rate Is our site the source? High presence, but others control the story
Sentiment How are we framed? Winning mentions that quietly hurt you

Layer two: what it drives

Presence is the leading indicator; this is the lagging one that earns the budget.

4. AI-referral revenue

Tie the AI traffic you can see to outcomes, not just sessions, but conversions and revenue. AI-referred visitors tend to arrive high-intent (they already got a recommendation), so this channel often punches above its session count. Reported as revenue, it's the line that turns "we're more visible" into "it's worth funding."

5. Dark-traffic index

The honest asterisk on metric 4. A large share of AI-driven visits arrive with no referrer and get filed as "Direct," so revenue measured on visible AI referrals understates the truth. A dark-traffic index. Estimated from deep-landing, new-user "Direct" sessions. Sizes the gap so you're not crediting AI with only the fraction you can see. (The full method is in why your AI traffic shows up as "Direct.")

Layer one tells you why the revenue is moving; layer two tells you whether it is. Report either alone and someone can wave it away. Report both and the channel defends itself.

How to read the stack together

The metrics are most useful as a diagnosis, not five separate gauges:

  • High SoV, low citation rate → you're known, but third parties own the narrative. Fix your own citable content.
  • High presence, poor sentiment → you appear in the wrong frame. This is a positioning and corroboration problem, not a coverage one.
  • Strong answer-level metrics, flat AI-referral revenue → either attribution is leaking into "Direct" (check your dark-traffic index) or the answers cite you without sending traffic, which is still brand impact in a zero-click world.

Keep your existing SEO and analytics. They still feed AI Overviews and still count real visits. But add this layer on top, sampled across engines and tracked as a trend. Standing up the answer-level three by hand is slow and non-reproducible; doing it continuously, across every engine, is exactly what Buffy Intel is built to report.

Frequently asked

What metrics should I track for AI visibility?

Five, in two layers. Answer-level (what the engines say): share of voice, citation rate, and sentiment. Sampled across many prompts and every engine. Outcome-level (what it drives): AI-referral revenue, and a dark-traffic index that estimates the AI visits hiding in your 'Direct' channel. Together they answer 'do we appear, are we the source, how are we framed, and does it pay?'

Isn't share of voice enough on its own?

No. Share of voice tells you how often you appear versus competitors, but not whether your own site is the cited source, how you're being described, or whether any of it converts. A brand can have high presence and still be framed as 'the budget option' or cited entirely through third-party pages it can't control. You need the full stack to act on it.