The fastest way to lose a budget for AI visibility is to report a number nobody trusts. "We're mentioned more!" invites the obvious question, so what?, and a single presence metric can't answer it. To survive a planning review, AI-visibility reporting needs to connect what the engines say to what it drives.
Here's the stack we'd defend: five metrics, two layers. The first three are answer-level: measured inside the AI answers themselves. The last two are outcome-level: measured in your own analytics.
Layer one: what the engines say
These can't be pulled from a rank tracker or GA4. You have to ask the engines the questions your customers ask, across many prompts and every engine, and log what comes back, because answers vary run-to-run, and they disagree with each other.
1. Share of voice
How often you appear, relative to competitors, across a representative set of buyer prompts. This is the headline presence number, but read it as share, not raw count: appearing in 40% of answers means little until you know a competitor is in 80%. Track it per engine, because ChatGPT, Gemini, Claude, and Google's AI surfaces each draw a different shortlist.
2. Citation rate
When you are mentioned, is your own site the cited source, or is the engine learning about you through a retailer, a review site, or a competitor's comparison page? High presence with low citation rate means the narrative about you is being written by others. This is the most actionable metric in the stack, because it points straight at content you can fix.
3. Sentiment
How you're described, not just whether you appear. "Premium and well-reviewed" and "a cheaper alternative" are both mentions; only one helps you. Sentiment turns presence into positioning, and it's where a brand most often discovers the gap between how it sees itself and how the models summarise it.
| Answer-level metric | Question it answers | The trap if you ignore it |
|---|---|---|
| Share of voice | Do we appear, vs competitors? | Counting raw mentions with no benchmark |
| Citation rate | Is our site the source? | High presence, but others control the story |
| Sentiment | How are we framed? | Winning mentions that quietly hurt you |
Layer two: what it drives
Presence is the leading indicator; this is the lagging one that earns the budget.
4. AI-referral revenue
Tie the AI traffic you can see to outcomes, not just sessions, but conversions and revenue. AI-referred visitors tend to arrive high-intent (they already got a recommendation), so this channel often punches above its session count. Reported as revenue, it's the line that turns "we're more visible" into "it's worth funding."
5. Dark-traffic index
The honest asterisk on metric 4. A large share of AI-driven visits arrive with no referrer and get filed as "Direct," so revenue measured on visible AI referrals understates the truth. A dark-traffic index. Estimated from deep-landing, new-user "Direct" sessions. Sizes the gap so you're not crediting AI with only the fraction you can see. (The full method is in why your AI traffic shows up as "Direct.")
Layer one tells you why the revenue is moving; layer two tells you whether it is. Report either alone and someone can wave it away. Report both and the channel defends itself.
How to read the stack together
The metrics are most useful as a diagnosis, not five separate gauges:
- High SoV, low citation rate → you're known, but third parties own the narrative. Fix your own citable content.
- High presence, poor sentiment → you appear in the wrong frame. This is a positioning and corroboration problem, not a coverage one.
- Strong answer-level metrics, flat AI-referral revenue → either attribution is leaking into "Direct" (check your dark-traffic index) or the answers cite you without sending traffic, which is still brand impact in a zero-click world.
Keep your existing SEO and analytics. They still feed AI Overviews and still count real visits. But add this layer on top, sampled across engines and tracked as a trend. Standing up the answer-level three by hand is slow and non-reproducible; doing it continuously, across every engine, is exactly what Buffy Intel is built to report.