Field note

What ChatGPT's query fan-out targets: the site: operator dataset (2026)

A July 2026 Peec AI analysis found ChatGPT 5.6 uses the site: search operator in about 43% of its query fan-outs (versus 0.004% in ChatGPT 5.5), and that 84% of those site: domains are the brand's own first-party site. Reddit accounts for 71.5% of the non-branded targets. This is a fully-attributed reference to what the fan-out actually points at, with every figure dated and hedged.

Buffy Editorial2026-07-24 · 6 min read

ChatGPT's query fan-out increasingly narrows its background searches to a single named domain using the site: operator, and the domains it points at are overwhelmingly first-party. According to a July 2026 analysis by AI-visibility platform Peec AI, the site: operator appeared in about 43% of ChatGPT 5.6 (codename Luna) query fan-outs, up from roughly 0.004% in ChatGPT 5.5, and 84% of those site: domains were branded — the site of the product or company the question was about. This reference lays out what the fan-out actually targets, with every figure attributed and dated.

Last reviewed: 24 July 2026. All figures below are from Peec AI's analysis of ChatGPT 5.6 query fan-outs, published by researcher David Konitzny (22 July 2026), unless noted. It is a single-vendor, directional dataset from one analytics vendor's prompt sample, so treat the exact percentages as a snapshot of one model version and cite "Peec AI, July 2026" with the date when you reuse a figure. For the underlying mechanics of the fan-out step itself, see how query fan-out works.

How much more does ChatGPT 5.6 use the site: operator?

Dramatically more than the prior version. The site: operator restricts a search to one domain, and its use inside ChatGPT's fan-out jumped by orders of magnitude between versions. Peec AI's term-frequency comparison of the sub-queries each model generates:

Term in fan-out sub-queries ChatGPT 5.6 (Luna) ChatGPT 5.5
site: prefix 43.2% 0.004%
Year mention (2025 / 2026) 26.5% 6.2%
"official" 21.9% 1.4%
"best" / "top" 5.4% 8.5%
"pricing" / "cost" 3.0% 0.51%
"review" 1.1% 1.02%
"vs" / "versus" 0.03% 1.30%

The pattern: Luna leans heavily on domain-restricted, "official," and recency-dated sub-queries, and away from open "best/top" and "vs" comparisons. The fan-out is behaving less like a broad keyword sweep and more like targeted verification against specific, current, first-party sources.

Which domains does the site: operator target?

Mostly one domain at a time, and mostly the brand's own. Peec AI reports that in 88.66% of fan-outs using the operator, only a single site: domain appears; two or three domains show up in roughly 9% of cases combined. And the domains split sharply between first-party and secondary sources:

Domain type Share of site: domains What it means
Branded (first-party) 83.63% The model targets the product/company's own site
Non-branded (secondary) 16.37% Community, review, and institutional sources

That 84%-branded finding is the mechanistic companion to ChatGPT's shift toward brand-name fan-out sub-queries: the model has often already decided which brand the question is about and points a site: probe straight at that brand's domain. The consequence for AI visibility is direct — if your first-party pages hide their facts behind JavaScript or images, the fan-out narrows to your domain and finds nothing citable.

What terms appear alongside the site: operator?

Product attributes lead. When Peec AI clustered the terms that appear with a site: domain, the most common were product specifications, followed by "official" and reviews:

Term cluster alongside site: Share of site: occurrences
Product specs (size, model, colour) 11.72%
Official 9.29%
Reviews 4.71%
Support 3.17%
Pricing 3.14%
Recommendations 0.88%

The read: the operator is used to pull highly specific facts from a trusted domain — the exact spec, the official detail, the price — not to browse. This is why structured, extractable product facts on your own pages matter so much for the branded majority of these searches.

Which non-branded domains does the fan-out reach for?

Reddit, by a wide margin. Within the 16% of site: domains that are non-branded secondary sources, one platform dominates:

Non-branded domain Share of non-branded site: targets
reddit.com 71.5%
trustpilot.com 5.92%
linkedin.com 2.34%
g2.com 2.27%
capterra.com 1.75%

Reddit alone is about 12.28% of all site: operator domains — roughly one in eight. Grouped by category, the non-branded targets break down as Social & forums 76.6%, Review & comparison 10.9%, and Public institutions 4.1%. Within Social & forums, Reddit is 93.4% of the category; within Review & comparison, Trustpilot leads at 54.5% (then G2 20.9%, Capterra 16.1%). Note the important nuance from the cross-engine citation data: being the top retrieval target is not the same as being cited — ChatGPT rejects the vast majority of Reddit pages it pulls in.

Public institutions are more spread out (bbb.org 34.6%, fda.gov 22.8%, mayoclinic.org 17.7%, pubmed 8.6%), with arXiv appearing at about 7.1% of that category — echoing the silent-source pattern where research repositories are retrieved for context without always being cited.

Does the operator target whole domains or specific pages?

Whole domains, mostly. Peec AI found the site: operator points at the bare root domain in 86.19% of cases and at a specific path or URL in only 13.81%. But the intent differs by depth:

Term cluster Root-domain site: Domain + path site:
Product / specs 12.37% 32.52%
Pricing 5.41% 12.31%
Official / verification 10.17% 8.58%
Reviews / ratings 7.27% 0.55%
Customer service 4.41% 0.63%
Comparison 1.84% 0.05%

Root-domain searches are used to identify a trusted source; path-level searches appear when the model already has a precise need, such as a specific product detail or price.

In plain terms: broad informational intents (reviews, official, service, comparison) skew to the whole domain, while product- and price-specific queries reach for a specific page. The fan-out is a structured retrieval strategy — first identify the trusted domain, then narrow to the exact fact.

How should you use this dataset?

As a directional reference, not a target to game:

  • Cite the source and date. Attribute figures to "Peec AI, July 2026 (ChatGPT 5.6 analysis)" and hedge them as single-vendor and version-specific — Luna's behaviour can change with the next model release.
  • Serve your first-party facts cleanly. With 84% of site: targets being branded, your own pages must expose specs, pricing, and official details as server-rendered text the model can lift.
  • Know your non-branded destinations. For your category, the external validation the fan-out reaches for likely lives on Reddit, Trustpilot, G2, or Capterra — but weight that against whether each engine actually cites those sources.
  • Measure retrieval and citation separately. Being targeted by a site: probe is upstream of being cited, which is upstream of being recommended.

Tracking what the fan-out points at for your buyers' questions — and whether those retrievals turn into citations across ChatGPT, Gemini, Claude, and Google — is exactly what Buffy Intel is built to measure, engine by engine, over time.

Frequently asked

How often does ChatGPT use the site: operator in its query fan-outs?

In about 43% of them, per a July 2026 Peec AI analysis of ChatGPT 5.6 (codename Luna) query fan-outs, up from roughly 0.004% in ChatGPT 5.5. In other words, the model now routinely restricts a sub-search to a single named domain rather than searching the open web. It is single-vendor, directional data from one analytics vendor's prompt sample, so read the direction as firmer than the exact percentage.

Which domains does ChatGPT target most with the site: operator?

Mostly first-party brand domains. Peec AI reports 84% of site: operator domains are branded, meaning the model points at the site of the product or company the question is about. Only 16% are non-branded secondary sources, and within that non-branded group Reddit dominates at about 71.5%, followed distantly by Trustpilot, LinkedIn, G2, and Capterra. Most fan-outs (about 89%) use the site: operator on just one domain.

What does the site: operator data mean for AI visibility?

It means the fan-out is a precise, source-targeted retrieval step, not a blind keyword search. Because 84% of site: targets are the brand's own domain, your first-party pages must expose the facts (specs, pricing, official details) as clean server-rendered text, or the model narrows to your domain and finds nothing citable. Because Reddit and Trustpilot dominate the non-branded targets, external validation for your category tends to be pulled from those specific platforms.