To vet an AI-visibility measurement provider, judge the method, not the dashboard — because you cannot check an AI answer against a "correct" one, only check how the provider produced its number. The IAB's August 2026 framework, Measuring Visibility in the AI Era, turns that into a concrete disclosure checklist. This walks through six questions to ask any vendor before you buy, each tied to a specific thing the framework says a provider should be able to tell you.
Two tools can report different share-of-voice figures for the same brand in the same category, and both can be internally consistent — the divergence lives in method. The framework's governing rule is worth keeping in front of you the whole way through: where a provider cannot or will not disclose a required item, that absence is itself the answer. This checklist is neutral; it applies to every tool, Buffy Intel included.
The six questions at a glance
| # | Ask the provider | What a good answer looks like |
|---|---|---|
| 1 | Which platforms and model versions do you cover? | Named engines + model versions; a substantial majority of consumer AI traffic |
| 2 | How many queries, and how often? | Large, category-plus-subcategory set; 50+ minimum; weekly for decision-grade |
| 3 | How is your prompt library built and sourced? | Disclosed intent mix; grounded in real search data, not only synthetic |
| 4 | How do you sample each query? | Many responses per query, reported as a range — not a single response |
| 5 | How do you handle hallucinations and errors? | Detected, reported per platform, surfaced — never silently filtered |
| 6 | How do you manage baselines when a model updates? | Documented re-baselining; platform-driven shifts separated from real ones |
1. Which AI platforms and model versions do you cover?
Start here, because coverage caps everything else. A tool that watches only one engine sees a fraction of the picture, since engines cite and recommend very different pages. The framework's decision-grade bar is coverage of platforms that collectively represent a substantial majority of consumer AI traffic in your market — with per-platform results shown, not blended into one figure that hides the divergence. Ask for model-version specificity too: "ChatGPT" is not an answer; the underlying model changes behaviour.
2. How many queries do you issue, and how often?
Volume and cadence are the clearest directional-vs-decision-grade line. The framework treats fewer than 50 queries as exploratory — not even directional — and wants a large, diverse set with subcategory coverage (an "office chairs" program should also probe "ergonomic office chairs" and "office chairs for back pain"). On cadence: monthly or quarterly sampling is fine for trend-watching; weekly or more frequent is expected before you move budget on the data. Match the cadence to the decision — measuring weekly for a quarterly decision just buys you noise.
3. How is your prompt library built and sourced?
This is the most consequential choice a provider makes and the least visible to buyers. A query set heavy on "best X" recommendation prompts flatters brands that do well in recommendations and understates them everywhere else. Ask the provider to disclose its intent-type mix across the four types the framework names — informational ("what is X"), comparison ("X vs Y"), recommendation ("best X for Y"), and transactional ("where to buy X") — and whether queries are grounded in real consumer search data or generated synthetically. As the framework warns, strong reproducibility on a badly built prompt library just gives you "precise answers to the wrong questions."
4. How do you sample each query — once, or many times?
This is the question that separates real measurement from a screenshot. Because AI answers are non-deterministic, the same query issued twice returns different brands, sources, and framing. The framework's line is that "single-response measurement is not measurement" — a brand's visibility on a query is a distribution, not a value. A credible provider samples each query multiple times and reports a range with a stated variability, so you can tell a real four-point move from ordinary jitter.
A four-point change in share of voice is only meaningful if it is larger than the variation you'd see from re-issuing the same queries against the same platforms on the same day. Without a variability baseline, trend interpretation is guesswork.
5. How do you handle hallucinated and factually inaccurate mentions?
An AI answer can invent an association your brand has no connection to, or attach the wrong price or spec to a real mention. Both distort your metrics and carry brand-safety risk. The framework requires providers to detect both, report them per platform, and surface flagged mentions to you rather than silently excluding them — because you have a legitimate interest in knowing when and how AI misrepresents you. Treat the absence of any documented hallucination-detection method as a material gap in a quality claim, not a rounding error.
6. How do you manage baselines when a model changes?
A model update can move your share of voice several points overnight with no change to your brand at all. The distinguishing skill of a serious program is telling a platform-driven shift (it appears across many brands or the whole category on one engine) from a market-driven one (isolated to your brand or specific query types). Ask how the provider documents re-baselining events — when one happened, what triggered it, and whether it reports pre- and post-update data as separate series rather than stitching them into one misleading trend line.
Match the disclosure tier to the decision
The framework's closing move is a two-tier disclosure model: minimum disclosure (platform coverage, collection architecture, summary-level prompt-library and attribution logic) is enough to place a tool's output in context for directional use; enhanced disclosure (full query sourcing, prompt-library detail, panel validity, factual-accuracy classification, historical versioning) is what you should expect from anything positioned as decision-grade. A provider's willingness to give enhanced disclosure is itself a signal of methodological confidence.
Buffy Intel is built to answer these six questions out of the box — multi-engine coverage, hundreds of prompts sampled repeatedly, results reported as trends with their variability, and hallucinated or wrong mentions flagged rather than hidden. See how Buffy Intel measures AI visibility, then hold us to the same checklist you'd hold anyone.