Field note

Which schema types does the web actually use? Google and Schema.org's usage dataset

In June 2026 Schema.org and Google published a public dataset of how often every Schema.org type appears across the crawlable web, bucketed by number of domains. This is a fully-sourced reference to what it contains, the domain buckets, the most-used types, and how to read it for AI visibility. Every figure attributed and dated.

Buffy Editorial2026-07-01 · 5 min read

In June 2026, Schema.org and Google published a public dataset showing how often each Schema.org type and property appears across the web Google crawls. It answers a question SEO and content teams have argued about for a decade, which structured-data types are actually used, and by how many sites, with crawl data instead of guesses. This is a dated reference to what the dataset contains, how to read its domain buckets, and what it does (and does not) mean for AI visibility.

Last reviewed: 1 July 2026. All figures below come from Schema.org's announcement (4 June 2026), its usage-statistics documentation, and the dataset's first release (dated May 2026), corroborated by Search Engine Land and industry coverage. The counts are Google's crawl-based measurements, aggregated at the domain level and reported as ranges, so read the direction as firmer than any single number, and cite "Schema.org/Google, May 2026 release" when you reuse a figure.

What is in the Schema.org usage dataset?

The dataset reports, for every Schema.org term, how many distinct domains use it. Schema.org announced it on 4 June 2026 as a collaboration with Google, framed as transparency into how the vocabulary is actually deployed on the web.

Attribute Detail
Announced 4 June 2026 (Schema.org blog)
First data release May 2026 (file 2026_05.csv)
Coverage 958 Itemtypes (types) + 4,587 predicates (properties) = 5,545 entries
Formats CSV and JSON
Location Canonical Schema.org GitHub repository (schemaorg/schemaorg, under the Google public-stats path)
Update cadence Monthly: a new file per month, so trends are trackable over time
Measurement Term frequencies within Google's public web crawl

Source: Schema.org (blog + usage-statistics docs, 2026); coverage counts via industry reporting (ppc.land). Because a new file lands each month, the repository doubles as a longitudinal record. You can diff 2026_05.csv against later months to watch adoption move.

How does the "domain bucket" work, and why ranges, not counts?

The dataset counts domains, not pages or markup objects. Schema.org states the rule plainly: "if you use the same term on 100 pages of your site, it still only counts as one domain using it." A publisher with Product markup on half a million URLs is one domain in the tally, exactly like a shop with three.

Instead of exact totals, each term is placed in a popularity range bucket. The buckets seen in the first release span several orders of magnitude:

Domain-count bucket Reading
10M+ Near-universal. Foundational site-structure and identity types
1M - 10M Very common. Mainstream commerce and content types
100K - 1M Widely used within their niche
10K - 100K Established but specialised
1K - 10K Niche / vertical-specific
< 1K Rare. Long-tail, experimental, or highly specialised terms

Ranges exist for privacy and stability: aggregating to buckets avoids exposing exact per-site data and, as Schema.org notes, filters daily noise because "web adoption trends change slowly." The practical effect is that the dataset tells you an order of magnitude. "Millions of sites" versus "a few thousand", not a precise count.

Which schema types does the most of the web use?

The top bucket is dominated by the types that describe site structure and identity, not fancy rich-result markup. Per the May 2026 release, the most widely adopted Itemtypes are:

  1. WebPage
  2. WebSite
  3. Organization
  4. BreadcrumbList
  5. ListItem

Commerce aggregate types sit a tier down. AggregateOffer and AggregateRating land around the 1M-10M-domain range, while specialised types such as AllocateAction, AgreeAction, and AmpStory fall into the under-1K long tail. The pattern is intuitive once stated: almost every site declares what page this is and who publishes it; far fewer mark up ratings, and only a handful use the exotic action types.

The web agrees on the boring basics. WebPage, WebSite, and Organization are near-universal; the specialised types most SEO advice obsesses over live in the long tail. Adoption tells you what is common, not what will get you cited.

The counts confirm the entity-clarity fundamentals we already recommend: Organization and WebSite markup. The schema that tells an engine who you are and what site this is. Is table stakes, used by the bulk of the crawlable web.

What does this dataset mean for AI visibility?

It is a prioritisation input, not a ranking factor. Popularity in the dataset measures adoption, not citation performance. A type used by ten million domains does not earn you a citation, and a rare type is not automatically valuable. What the dataset usefully does:

  • Settles "is this type real / worth it?" debates. Knowing a type sits in the 10M+ bucket versus the under-1K tail helps you decide whether it is a safe, well-supported choice or an experimental one.
  • Reveals table-stakes markup. If the near-universal types (WebPage, WebSite, Organization, BreadcrumbList) are missing from your pages, you are below the web's baseline for machine-legibility. See making your site agent-readable.
  • Guides which valid types to add next, which we turn into a step-by-step method in how to prioritise your structured data using this data.

The caveats matter, and Schema.org states them. The counts reflect the web as indexed by Google, so sites that block crawlers are undercounted. Markup formats. JSON-LD, Microdata, RDFa: are combined into one figure. And niche terms (medical, government) stay in low buckets despite being exactly right for their domain. The honest reading: use adoption data to avoid dead-end markup and to confirm the basics, but never treat "popular" as "better for citations."

Structured data still earns its place because it makes your facts easy for an engine to lift. The extractability half of getting cited by AI and of accessibility-as-parseability. For commerce specifically, Product/Offer markup is what lets an engine quote your price and availability cleanly, a point we develop in writing conversational product pages. Schema is necessary, not sufficient: the markup has to match real content, and the facts underneath still have to be specific, accurate, and corroborated.

Because the dataset refreshes monthly and adoption drifts, we keep references like this on a refresh cadence and update the buckets substantively as new files land. Never a silent date-bump. The number that ultimately matters is not how many sites use a schema type, but whether the engines surface, cite, and recommend your brand for the questions your buyers actually ask. Tracking that across every engine, over time, is exactly what Buffy Intel is built to do.

Frequently asked

What is the Schema.org usage statistics dataset?

It is a public dataset, announced by Schema.org and Google on 4 June 2026, that reports how many domains use each Schema.org type and property across the web Google crawls. Counts are aggregated at the domain level and shown as popularity ranges (buckets like 10M+, 1M-10M, down to under 1K) rather than exact figures. The first release, dated May 2026, covers 958 Itemtypes and 4,587 predicates, and it is published as CSV and JSON on the Schema.org GitHub repository, updated monthly.

Which schema types are used by the most websites?

In the first release, the top bucket (roughly 10M+ domains) is led by the site-structure and identity types: WebPage, WebSite, Organization, BreadcrumbList, and ListItem. Commerce aggregate types such as AggregateOffer and AggregateRating sit a tier lower, around the 1M-10M-domain range. These figures are Google's crawl-based counts as of the May 2026 release; treat the ranking as directional, not exact.

Does using a popular schema type improve AI visibility?

Not on its own. Popularity in the dataset tells you what is widely adopted, not what earns citations. Structured data helps AI parse and lift your facts cleanly, but the schema has to match real page content, and the underlying facts still have to be specific, accurate, and corroborated. Use the dataset to prioritise which valid types to implement, then measure whether adding them actually moves your citations.