DATAGEO · Jun 16, 2026 · 5 min read

Schema.org Just Published Its First Usage Dataset — Here's What It Reveals

// TL;DR

On June 4, 2026, schema.org published the first public dataset showing which Schema.org types are actually deployed across the web. The data is authoritative, domain-level, and updated monthly. What it reveals: structured data adoption is still low enough that deploying the right types—especially the ones AI systems use for citations—is a genuine competitive moat. Most businesses are missing the exact schema types that drive AI visibility.

Most businesses think structured data is either too complex or too niche to matter. The data says otherwise.

On June 4, 2026, schema.org—in collaboration with Google—published something the industry has needed for years: a monthly Usage Statistics Dataset showing which Schema.org Types and Properties are actually deployed across the public web.

This is the first time we've had authoritative, aggregate, domain-level visibility into real-world structured data adoption at scale. And what it reveals is both opportunity and warning.

What the Dataset Actually Is

The Usage Statistics Dataset is aggregate data pulled from Google's web index, tracking which Schema.org types and properties appear on domains across the public web. It's domain-level (not page-level), monthly-updated, and publicly available.

For the first time, you can see:

  • Which schema types are most commonly deployed (e.g., Organization, Article, Product)
  • Adoption rates by type—what percentage of domains use each schema
  • Property-level usage within each type (which fields businesses are actually filling out)
  • Trend data month-over-month as adoption shifts

This isn't survey data or scraper estimates. It's Google's own index. If you want to know what structured data the web is actually running—not what SEO blogs recommend—this is it.

The Adoption Gap: "Recommended" vs. Deployed

The first thing the data reveals is how wide the gap is between what's recommended and what's actually deployed.

SEO documentation has recommended schema types for years—FAQPage for FAQ sections, HowTo for tutorials, Product for e-commerce, LocalBusiness for brick-and-mortar. But the usage data shows that adoption is still shockingly low, even for types that have been standard practice since 2018.

Why does this matter? Because AI systems—ChatGPT, Claude, Perplexity, Google AI Overviews—are trained on and retrieve from structured data signals. If your schema types match what's in the training corpus and the retrieval index, you're more likely to be cited.

Most businesses assume structured data is table stakes. The usage dataset proves it's still an edge. And edges compound.

Which Schema Types AI Systems Actually Use

Not all schema types are created equal when it comes to AI citations. Some types appear disproportionately in AI-surfaced content because they map directly to the questions users ask.

The types that matter most for GEO:

Article — AI systems pull from articles for explanatory, educational, and how-to queries. If your blog content doesn't have Article schema with headline, author, datePublished, and publisher fields, you're invisible in knowledge-retrieval queries.

Organization — Brand-authority queries (“who is X,” “what does Y do”) pull heavily from Organization schema. This is the baseline schema every business should deploy. Most don't.

LocalBusiness — Local queries (“dentist near me,” “best coffee shop in Brooklyn”) rely on LocalBusiness schema to surface candidates. AI Overviews and Perplexity's local results are schema-first—no schema, no citation.

FAQPage / HowTo — Instructional and procedural queries (“how to do X,” “why does Y happen”) preferentially cite FAQPage and HowTo markup because it's structured for direct-answer extraction. If your content answers questions but doesn't have the schema, AI systems skip it.

Product — E-commerce and product-research queries pull from Product schema. If your product pages lack structured pricing, availability, and review data, you're not entering the consideration set when users ask AI to compare options.

The usage dataset shows these five types have the highest deployment rates—but “highest” is relative. Adoption is still low enough that deploying them correctly puts you ahead of most competitors.

Why This Data Changes the Competitive Landscape

For years, structured data has been in the “SEO best practices” bucket—something technical teams should do eventually, when they get around to it. The usage dataset changes that framing.

Now you can see exactly which schema types your competitors are deploying and where the gaps are. You can benchmark your own coverage against industry averages. You can track month-over-month adoption trends and predict when a type will shift from “edge” to “table stakes.”

More importantly, you can prioritize which types to deploy based on real-world data, not SEO folklore. If a schema type is heavily used in AI citations but under-deployed across the web, that's a high-ROI target.

The businesses that treat this dataset as competitive intelligence will build citation moats. The ones that ignore it will wonder why their competitors are being cited and they're not.

The Tactical Play: Deploy Now, Before It's Table Stakes

Structured data adoption follows a predictable curve. Early adopters get disproportionate visibility. Late adopters get parity at best.

Right now, the usage data shows we're still early enough that deploying the right schema types is a genuine advantage. But that window is closing. As more businesses wake up to AI citability and start instrumenting their sites with structured data, the relative edge shrinks.

Here's the move: Audit your current schema coverage. Compare it against the types AI systems use for citations (Article, Organization, LocalBusiness, FAQPage, HowTo, Product). Deploy the gaps. Use the usage dataset to benchmark yourself against industry norms and identify under-deployed types where you can over-index.

This isn't “nice to have” SEO work anymore. It's competitive positioning. The businesses that ship schema now—while adoption is still low—will own AI citation share for the next two years. The ones that wait will spend that time trying to catch up.

The Bottom Line

The Schema.org Usage Statistics Dataset is the first authoritative view into what structured data the web is actually running. And what it shows is clear: adoption is still low enough that getting it right is a moat.

AI systems cite structured data. The types they use most—Article, Organization, LocalBusiness, FAQPage, HowTo, Product—are still under-deployed. The businesses that close that gap now will win AI visibility for years. The ones that assume “everyone does this already” will lose.

The data is public. The opportunity is real. The window won't stay open.

// want the full picture?

We run GEO audits that score how visible your business is to AI systems—and show you exactly where you're being left out of the answer.

get your GEO audit →