agentrichy · research
Field scan · 133,000 domains · 3 sources

Is the web ready for AI agents?

We pointed our own scan engine at the open web — the way an AI crawler actually fetches, with no JavaScript — across three independent domain rankings. The head of the web is tidier than the tail. It is not much more ready.

The state of play — top 100,000 sites

Even the most popular sites leave AI reading between the lines

Across the 32,846 top-ranked sites that returned a clean page, the median AI-readiness grade is a C. The gaps aren't exotic — they're the basics that make a page machine-legible.

74/100
Median readiness score — grade C
35%
Have valid structured data (JSON-LD)
18%
JavaScript-gated — invisible to no-JS crawlers
38%
Ship no sitemap at all
Triangulation — why believe this

Three independent rankings, one answer

Median score across four samples drawn from three ranking methodologies (traffic-aggregated, backlink-based, PageRank). The two popularity lists agree almost exactly — so the "head" numbers are real, not one list's artifact. And the true long tail scores only a few points lower.

Median AI-readiness score, by source
clean-scanned sites per sample shown in tooltip
Show data table
SampleMedianValid JSON-LDBlocks AI-searchN
Majestic (backlinks)7540.2%17.5%1,409
Tranco top-100k (head)7435.0%13.6%32,846
DomCop long-tail (rank >1M)7131.7%7.8%2,249
Tranco random top-1M7029.0%18.2%1,880
The long tail is barely worse — but blocks AI half as often.

Obscure sites (rank >1M) score a median 71 — modern site builders auto-generate decent markup even for nobodies. Yet they block AI-search crawlers only 7.8% of the time, versus ~14–18% for popular sites. AI-blocking is a popular-site behavior; the long tail leaves the door open.

Distribution

The head of the web outscores the tail — modestly

Grades for the top 100,000 sites versus the true long tail (DomCop rank >1M). Popularity shifts the curve right, but the modal site is still a C, and roughly one in five still fails outright.

AI-readiness grade distribution
share of scanned sites, by grade
Show data table
GradeTop 100kLong tail
A (90–100)11.0%10.9%
B (80–89)24.2%21.0%
C (70–79)24.3%20.9%
D (60–69)18.8%22.1%
F (<60)21.7%25.0%
Head vs tail

Popularity buys tidiness, not readiness

Each metric plotted for the true long tail and the top 100,000. Schema and meta improve most with popularity. The sitemap gap barely moves — a blind spot the biggest sites share with everyone else.

Where popularity helps — and where it doesn't
% of scanned sites passing, long tail → head
Show data table
MetricLong tailTop 100k
Median score7174
Server-rendered content73.6%72.9%
Clean meta/semantics44.0%56.0%
Any JSON-LD37.4%41.2%
Valid JSON-LD31.7%35.0%
Has a sitemap60.0%61.9%
Check by check — top 100,000

Structured data is the web's biggest AI blind spot

Every check, sorted worst-first. Machine-readable data and content negotiation are emerging signals almost nobody ships yet (a neutral "warn"); structured data is the one with a hard, avoidable fail on 58% of top sites.

Per-check outcome, top 100,000 sites
100% of scanned sites · pass / warn / fail
Show data table
CheckPassWarnFail

Loading speed and llms.txt are excluded: speed was measured synthetically from one location (an artifact), and the llms.txt check over-counts sites that soft-404. Both are being corrected.

By top-level domain

.edu and .ai lead; national domains split on blocking

Mean readiness score by TLD (top 100k, TLDs with 150+ clean scans). Universities (.edu) and AI-native domains (.ai) top the table — and .edu blocks AI crawlers least of all. Hover any bar for its structured-data and AI-blocking rates.

Mean AI-readiness score by TLD
top 100k · bar = mean score · hover for JSON-LD % and AI-block %
Show data table
TLDMeanValid JSON-LDBlocks AI-searchN
.co.uk blocks AI-search crawlers 25% of the time — .ru and .edu under 6%.

National web cultures diverge sharply: UK sites (news-heavy) fence AI out; Russian and educational domains largely leave it open. The same content type, different posture toward machine readers.

Access — who gets let in

Sites fence off AI training, but leave the door open to citation

Share of top-100k sites that block each AI crawler in robots.txt. The pattern is deliberate: model-training crawlers are blocked more than answer-search crawlers. Owners are saying "don't train on me — but do cite me."

Most-blocked AI crawlers
% of scanned top-100k sites blocking each, in robots.txt
Show data table
CrawlerPurposeBlocked
GPTBot (training) blocked ~2.4× as often as ChatGPT-User (search/browse).

14.2% block OpenAI's training crawler; 6.0% block its user-triggered fetcher. The web is drawing a line between being learned from and being recommended.

Who's missing — the blockers

The sites with the best content block the hardest

First, is our scanned set just the obscure end of the list? No. The 32,846 sites we read have a median rank of ~48,000 of 100,000 — statistically the same as the sites that failed. Clean-scan rate is nearly flat across popularity (44% at the very top, ~30% deeper), so the head isn't systematically excluded. The sample is a representative mix, not a tail.

Clean-scan rate by popularity rank
% of each 10k rank band that returned a readable homepage
Show data table
Rank bandClean-scan %Hard-block (403) %
1–10k44.0%9.9%
10–20k34.5%7.5%
20–30k30.4%8.0%
30–40k29.0%7.4%
40–50k32.4%9.1%
50–60k28.2%8.4%
60–70k32.0%8.6%
70–80k33.1%8.1%
80–90k35.2%8.4%
90–100k29.6%7.4%

What is excluded is deliberate: 8.3% hard-blocked our scanner outright (HTTP 403), and ~56% more failed on challenge pages or connection resets — the same walls an AI crawler hits. And the 403 list is a who's-who of the most valuable content on the web:

The AI labs, publishers, and knowledge sites block hardest.

Among the most-popular domains that returned 403 to an automated fetch: OpenAI & ChatGPT, Anthropic (claude.ai), Bloomberg, the Financial Times, Medium, Quora, Stack Overflow, Britannica, Fandom, ScienceDirect, ResearchGate, Cambridge, eBay, Etsy, Indeed, Trustpilot, Tripadvisor, Yelp, NIH.gov, the Library of Congress. The sites with the most quotable, authoritative content are the least readable to machines — so the citation and structured-data gaps are worse than any figure here, for exactly the content AI most wants to use.

8.3%
Hard-blocked our scanner outright (HTTP 403)
~48k
Median rank of scanned sites — a mix, not the tail
56%
Failed on challenge pages / connection resets
50%
Even the top 1,000 sites scanned clean only half the time

Does the IP we scan from skew this? We re-scanned 1,500 residential-clean sites from a data-center IP — the kind of network a real AI crawler actually uses. 95% behaved identically; only ~5% blocked the DC IP specifically (Condé Nast titles like Vanity Fair and GQ, a few universities, active data-center-range 403s). So the origin IP is a minor factor — and a genuinely DC-originated AI crawler would see slightly more blocking than reported here, never less.

*Not a reachability census — "failed" is dominated by bot-defense challenge pages and connection drops, not dead domains.

What kind of content blocks

The AI companies block scrapers the most

Taking a run of well-known domains per category and checking who let an automated reader in, one pattern is impossible to miss: the firms building AI are the most defended, and so is paywalled scholarship. General news, by contrast, mostly stays open.

7 of 15 leading AI companies hard-block an automated reader.

OpenAI, ChatGPT, Anthropic's claude.ai, Perplexity, Midjourney, Character.AI and Poe all return 403 to a plain fetch. The companies whose products depend on crawling the web are the least willing to be crawled themselves.

Who blocks, by content type
illustrative run of well-known domains per category · "blocked" = 403 or failed fetch
Show data table
CategoryCheckedBlocked/failedNotable hard-403
AI companies159OpenAI, claude.ai, Perplexity, Midjourney, Poe
Academic / reference128ScienceDirect, ResearchGate, Cambridge, Britannica, NIH.gov
Retail / marketplace126eBay, Etsy, BestBuy, Tripadvisor, Yelp
Social / Q&A125Quora, Stack Overflow, Medium, Fandom
News / publishers145Bloomberg, FT, Economist

The takeaways: paywalled scholarship (8 of 12 academic sites blocked — ScienceDirect, Cambridge, Britannica, NIH) and community knowledge (Quora, Stack Overflow, Medium — the Q&A the AI-licensing wars are fought over) are walled off; but most general newsrooms — NYT, BBC, the Guardian, the Washington Post, AP — stay open, and only premium financial titles (Bloomberg, FT, Economist) shut the door. By country domain, hard-403 rates run highest on .com.au (17%), .ua (17%), .ca (15%) and .com.br (14%) — lowest on .com (9%).

How to read this

Method & honest limits

Four samples across three independent rankings: the top 100,000 and a random 5,000 from Tranco (aggregated traffic); a random 3,000 from Majestic Million (backlinks); and a 5,000 long-tail sample from DomCop's Open PageRank 10M (ranks >1M). Each scanned with our production engine, no JavaScript, as a plain AI crawler fetches. Analysis is over the 38,400 sites that returned a clean 2xx homepage. Full source URLs, sampling commands, and citations are in the companion sources document.