Even the most popular sites leave AI reading between the lines
Across the 32,846 top-ranked sites that returned a clean page, the median AI-readiness grade is a C. The gaps aren't exotic — they're the basics that make a page machine-legible.
Three independent rankings, one answer
Median score across four samples drawn from three ranking methodologies (traffic-aggregated, backlink-based, PageRank). The two popularity lists agree almost exactly — so the "head" numbers are real, not one list's artifact. And the true long tail scores only a few points lower.
Show data table
| Sample | Median | Valid JSON-LD | Blocks AI-search | N |
|---|---|---|---|---|
| Majestic (backlinks) | 75 | 40.2% | 17.5% | 1,409 |
| Tranco top-100k (head) | 74 | 35.0% | 13.6% | 32,846 |
| DomCop long-tail (rank >1M) | 71 | 31.7% | 7.8% | 2,249 |
| Tranco random top-1M | 70 | 29.0% | 18.2% | 1,880 |
Obscure sites (rank >1M) score a median 71 — modern site builders auto-generate decent markup even for nobodies. Yet they block AI-search crawlers only 7.8% of the time, versus ~14–18% for popular sites. AI-blocking is a popular-site behavior; the long tail leaves the door open.
The head of the web outscores the tail — modestly
Grades for the top 100,000 sites versus the true long tail (DomCop rank >1M). Popularity shifts the curve right, but the modal site is still a C, and roughly one in five still fails outright.
Show data table
| Grade | Top 100k | Long tail |
|---|---|---|
| A (90–100) | 11.0% | 10.9% |
| B (80–89) | 24.2% | 21.0% |
| C (70–79) | 24.3% | 20.9% |
| D (60–69) | 18.8% | 22.1% |
| F (<60) | 21.7% | 25.0% |
Popularity buys tidiness, not readiness
Each metric plotted for the true long tail and the top 100,000. Schema and meta improve most with popularity. The sitemap gap barely moves — a blind spot the biggest sites share with everyone else.
Show data table
| Metric | Long tail | Top 100k |
|---|---|---|
| Median score | 71 | 74 |
| Server-rendered content | 73.6% | 72.9% |
| Clean meta/semantics | 44.0% | 56.0% |
| Any JSON-LD | 37.4% | 41.2% |
| Valid JSON-LD | 31.7% | 35.0% |
| Has a sitemap | 60.0% | 61.9% |
Structured data is the web's biggest AI blind spot
Every check, sorted worst-first. Machine-readable data and content negotiation are emerging signals almost nobody ships yet (a neutral "warn"); structured data is the one with a hard, avoidable fail on 58% of top sites.
Show data table
| Check | Pass | Warn | Fail |
|---|
Loading speed and llms.txt are excluded: speed was measured synthetically from one location (an artifact), and the llms.txt check over-counts sites that soft-404. Both are being corrected.
.edu and .ai lead; national domains split on blocking
Mean readiness score by TLD (top 100k, TLDs with 150+ clean scans). Universities (.edu) and AI-native domains (.ai) top the table — and .edu blocks AI crawlers least of all. Hover any bar for its structured-data and AI-blocking rates.
Show data table
| TLD | Mean | Valid JSON-LD | Blocks AI-search | N |
|---|
National web cultures diverge sharply: UK sites (news-heavy) fence AI out; Russian and educational domains largely leave it open. The same content type, different posture toward machine readers.
Sites fence off AI training, but leave the door open to citation
Share of top-100k sites that block each AI crawler in robots.txt. The pattern is deliberate: model-training crawlers are blocked more than answer-search crawlers. Owners are saying "don't train on me — but do cite me."
Show data table
| Crawler | Purpose | Blocked |
|---|
14.2% block OpenAI's training crawler; 6.0% block its user-triggered fetcher. The web is drawing a line between being learned from and being recommended.
The sites with the best content block the hardest
First, is our scanned set just the obscure end of the list? No. The 32,846 sites we read have a median rank of ~48,000 of 100,000 — statistically the same as the sites that failed. Clean-scan rate is nearly flat across popularity (44% at the very top, ~30% deeper), so the head isn't systematically excluded. The sample is a representative mix, not a tail.
Show data table
| Rank band | Clean-scan % | Hard-block (403) % |
|---|---|---|
| 1–10k | 44.0% | 9.9% |
| 10–20k | 34.5% | 7.5% |
| 20–30k | 30.4% | 8.0% |
| 30–40k | 29.0% | 7.4% |
| 40–50k | 32.4% | 9.1% |
| 50–60k | 28.2% | 8.4% |
| 60–70k | 32.0% | 8.6% |
| 70–80k | 33.1% | 8.1% |
| 80–90k | 35.2% | 8.4% |
| 90–100k | 29.6% | 7.4% |
What is excluded is deliberate: 8.3% hard-blocked our scanner outright (HTTP 403), and ~56% more failed on challenge pages or connection resets — the same walls an AI crawler hits. And the 403 list is a who's-who of the most valuable content on the web:
Among the most-popular domains that returned 403 to an automated fetch: OpenAI & ChatGPT, Anthropic (claude.ai), Bloomberg, the Financial Times, Medium, Quora, Stack Overflow, Britannica, Fandom, ScienceDirect, ResearchGate, Cambridge, eBay, Etsy, Indeed, Trustpilot, Tripadvisor, Yelp, NIH.gov, the Library of Congress. The sites with the most quotable, authoritative content are the least readable to machines — so the citation and structured-data gaps are worse than any figure here, for exactly the content AI most wants to use.
Does the IP we scan from skew this? We re-scanned 1,500 residential-clean sites from a data-center IP — the kind of network a real AI crawler actually uses. 95% behaved identically; only ~5% blocked the DC IP specifically (Condé Nast titles like Vanity Fair and GQ, a few universities, active data-center-range 403s). So the origin IP is a minor factor — and a genuinely DC-originated AI crawler would see slightly more blocking than reported here, never less.
*Not a reachability census — "failed" is dominated by bot-defense challenge pages and connection drops, not dead domains.
The AI companies block scrapers the most
Taking a run of well-known domains per category and checking who let an automated reader in, one pattern is impossible to miss: the firms building AI are the most defended, and so is paywalled scholarship. General news, by contrast, mostly stays open.
OpenAI, ChatGPT, Anthropic's claude.ai, Perplexity, Midjourney, Character.AI and Poe all return 403 to a plain fetch. The companies whose products depend on crawling the web are the least willing to be crawled themselves.
Show data table
| Category | Checked | Blocked/failed | Notable hard-403 |
|---|---|---|---|
| AI companies | 15 | 9 | OpenAI, claude.ai, Perplexity, Midjourney, Poe |
| Academic / reference | 12 | 8 | ScienceDirect, ResearchGate, Cambridge, Britannica, NIH.gov |
| Retail / marketplace | 12 | 6 | eBay, Etsy, BestBuy, Tripadvisor, Yelp |
| Social / Q&A | 12 | 5 | Quora, Stack Overflow, Medium, Fandom |
| News / publishers | 14 | 5 | Bloomberg, FT, Economist |
The takeaways: paywalled scholarship (8 of 12 academic sites blocked — ScienceDirect, Cambridge, Britannica, NIH) and community knowledge (Quora, Stack Overflow, Medium — the Q&A the AI-licensing wars are fought over) are walled off; but most general newsrooms — NYT, BBC, the Guardian, the Washington Post, AP — stay open, and only premium financial titles (Bloomberg, FT, Economist) shut the door. By country domain, hard-403 rates run highest on .com.au (17%), .ua (17%), .ca (15%) and .com.br (14%) — lowest on .com (9%).
Method & honest limits
Four samples across three independent rankings: the top 100,000 and a random 5,000 from Tranco (aggregated traffic); a random 3,000 from Majestic Million (backlinks); and a 5,000 long-tail sample from DomCop's Open PageRank 10M (ranks >1M). Each scanned with our production engine, no JavaScript, as a plain AI crawler fetches. Analysis is over the 38,400 sites that returned a clean 2xx homepage. Full source URLs, sampling commands, and citations are in the companion sources document.
- Every percentage is "of sites we could read," not of all sites. Of 100,000 top domains only ~33,000 returned a clean homepage; the rest — bot walls, HTTP 403/429, timeouts, dead domains — are excluded from the base. Because the fiercest AI-blockers are exactly the sites that fail the scan, blocking and quality figures are optimistic floors: the real web is a little worse, never better.
- Speed is excluded — synthetic, single-location, no field data.
- llms.txt is excluded — the current check over-counts soft-404s; true adoption is lower, pending a fix.
- Even the "random" popular samples are top-1M; only the DomCop long-tail escapes popularity bias. A uniform sample of all registered domains (zone files / CT logs) is future work.