AI search term

Holdout test

A measurement method that leaves a comparable group deliberately unchanged so a treated group's outcome can be read against a real counterfactual, isolating cause from correlation.

Also known as: holdout group, holdback test, hold-out experiment, control holdout, holdout experiment

Updated 2026-08-07

A holdout test is a way to measure whether something caused a result by keeping a comparable group unchanged for comparison. The changed group is the treatment; the untouched group is the holdout, or control. Because both groups drift together over time, any extra movement in the treated group, measured against the holdout, is attributable to the change rather than to background noise.

In AI visibility the holdout is the fix for a stubborn problem: citations rise and fall on their own, so a bare before-and-after cannot tell you whether your edit or the tide moved them. Hold out a comparable set of pages when you add schema, refresh copy, or build links, then read the treated set's citation coverage against the holdout, not against zero. The strongest version randomizes assignment: the 2026 Agarwal and Sen field experiment hid Google's AI Overview at random for some searches and compared clicks against a control, which is what let it claim a causal effect rather than a correlation.

Holdout tests have limits. The groups must be genuinely comparable, the window must outlast the crawl-to-cite lag, and on third-party surfaces you often cannot randomize at all, only approximate a control with matched pages or queries. Used honestly, though, a holdout is the difference between "cited pages have X" and "adding X made pages cited," and it is the discipline behind reading any AI-visibility claim, or measuring your own share of voice, without fooling yourself.