Q4 2026 Report

10.2% adoption

AI Discovery File Adoption Research

Measuring how the world's top websites prepare for AI systems. Based on a crawl of 1,995 domains.

1,896
Domains Crawled
of 1,995 total
10.2%
ADF Adoption
194 domains
2.2
Avg Readiness
out of 5.0
84.2%
No AI Policy
in robots.txt
You are viewing the archived Q4 2026 report. View latest report →

Summary

This Q4 2026 report analyses AI Discovery File adoption across 1,896 of the web's most prominent domains. 10.2% of domains have at least one AI Discovery File, 84.2% have no AI-specific crawler policy in their robots.txt, and the average AI readiness tier is 2.2 out of 5.0. Data is collected quarterly using the methodology described in our full methodology documentation.

AI Discovery Files are a set of 10 standardised root-level files, including llms.txt, ai.txt, ai.json, identity.json, and brand.txt, that help AI systems such as ChatGPT, Claude, and Gemini discover, interpret, and correctly represent a website. This research tracks their real-world adoption among the domains most likely to be referenced by AI systems when answering user questions.

Key Findings

The Q4 2026 crawl puts AI Discovery File adoption above 10% for the first time: 194 of 1,896 successfully crawled domains (10.2%) now publish at least one file, up from 164 (9.4%) in Q3. llms.txt still carries almost all of it, but the files it carries are getting better, and the number of AI-Ready sites rose by more than a third. This quarter is also the cleanest comparison so far. Only 99 domains errored, against 251 in July, so the sample is 152 domains larger and much closer to the full list. That cuts both ways: the percentages are more trustworthy, and some of the new adopters may simply be sites we could not reach last time rather than sites that changed.

  1. Adoption passes 10%, and the better crawl makes the rise easier to trust Domains with any AI Discovery File rose from 164 to 194 (+18%), and the share from 9.4% to 10.2%. The share matters more than the count here, because the crawl reached 1,896 domains this quarter against 1,744 in July (errors fell from 251 to 99). A rise in the percentage on a larger, more complete sample is the opposite of last quarter, when a shrinking sample flattered some numbers. We still cannot tell from aggregate data alone how many of the 30 new adopters were unreachable in July rather than newly published.
  2. llms.txt grows, and its quality improves faster than its count llms.txt was found on 150 domains, up from 122 (+23%), while valid and complete files rose from 82 to 110 (+34%). Invalid llms.txt files stayed exactly flat at 40, so every additional file this quarter was a good one: the valid share rose from 67% to 73%. Across all file types the count of complete files went from 85 to 110. Publishers are not only adopting llms.txt, they are getting the format right more often.
  3. AI-Ready sites up by more than a third, with Calendly, Check Point and Exeter reaching it for the first time Tier 4 (AI-Ready) domains rose from 44 to 61 (+39%), now 3.2% of the sample against 2.5% in Q3. Calendly, Check Point, Branch, hCaptcha and the University of Exeter reach AI-Ready for the first time in our data, and Dell and Greenpeace UK, both AI-Ready in Q2, are back after missing the July list. Movement runs both ways: Bunkbedsstore and Groupon UK still publish a valid llms.txt but fell to Tier 3, and Bluehost now returns 403 to our crawler while serving the same file to GPTBot and ClaudeBot (see the methodology note on user-agent gating). Tier 5 (AI-Optimised) is still empty after four quarters, and the average readiness score held at 2.2 out of 5.0.
  4. Blanket AI blocking falls, explicit permission rises again Sites blocking every AI crawler fell from 23 to 17 (1.3% to 0.9%), the first real drop in four quarters, while sites explicitly allowing AI crawlers rose from 31 to 38 (2.0%). Selective blocking edged up from 217 to 239 domains (12.6%), so the trend is towards deliberate, per-crawler decisions rather than all-or-nothing. The large majority is unchanged: 84.2% of sites still have no AI policy at all, the same share as in July.
  5. Other file types appear for the first time, but none of them is valid yet identity.json, brand.txt and developer-ai.txt were each found on a domain for the first time, and ai.txt and robots-ai.txt each on two. None of those new files passed validation, so they are attempts rather than adoption. llms.html moved the wrong way on quality: found on 49 domains (from 46), but valid files fell from 9 to 7 and complete files from 3 to 0. Breadth is starting, but only llms.txt is being implemented to a standard a machine can rely on.
  6. Still zero of 194 adopters reference a formal specification For the fourth quarter running, not one domain publishing an AI Discovery File references a formal specification. For scale, AI Discovery Files are now on 10.2% of crawled sites against 14.5% for security.txt and 2.3% for humans.txt, so this is no longer a fringe practice, yet every one of these files is still written without a shared schema or version to validate against. Adoption keeps outrunning standardisation, and the gap widens as adoption grows.

AI Visibility Research, October 2026

Changes from Q3 2026

Quarter-over-quarter changes in key metrics between Q3 2026 and Q4 2026.

ADF Adoption
9.4% → 10.2%
+0.8
Avg Readiness Score
2.2 → 2.2
0.0
Domains Crawled
1,744 → 1,896
+152
No AI Policy
84.2% → 84.2%
0.0pp
Per-file adoption change: Q3 2026 to Q4 2026
File Q3 2026 Q4 2026 Change
llms.txt 7.0% 7.9% +0.9
llms.html 2.6% 2.6% 0.0
ai.txt 0.1% 0.1% 0.0
ai.json 0.1% 0.1% 0.0
identity.json 0.0% 0.1% +0.1
brand.txt 0.0% 0.1% +0.1
faq-ai.txt 0.0% 0.0% 0.0
developer-ai.txt 0.0% 0.1% +0.1
robots-ai.txt 0.1% 0.1% 0.0

New to This Quarter's Top 20

  • branch.io
  • calendly.com
  • checkpoint.com
  • dell.com
  • exeter.ac.uk
  • greenpeace.org.uk
  • hcaptcha.com

Outside This Quarter's Top 20

  • bluehost.com
  • bunkbedsstore.uk
  • groupon.co.uk
  • hostgator.com.br
  • kingsfund.org.uk
  • klaviyo.com
  • life360.com

Ranked by readiness tier, then valid files. Most AI-Ready sites tie and ties fall alphabetically, so a site here may still be AI-Ready, and a site in the other list may be returning rather than new.

ADF Adoption by File Type

How many of the top websites have each AI Discovery File: and how many of those files pass structural validation. Files are checked at their canonical root-level URL (e.g., example.com/llms.txt) and validated against the ADF specification.

View data table
AI Discovery File adoption across 1,896 domains
File Found Valid Complete
llms.txt 150 110 110
llms.html 49 7 0
ai.txt 2 0 0
ai.json 2 0 0
identity.json 1 0 0
brand.txt 1 0 0
faq-ai.txt 0 0 0
developer-ai.txt 1 0 0
robots-ai.txt 2 1 0

AI Crawler Access Policies

How websites use robots.txt to manage access for 15 known AI user agents: from OpenAI's GPTBot to Anthropic's ClaudeBot. Each domain is classified into one of five access policies based on its aggregate behaviour across all agents. The per-agent table below shows which AI crawlers are most frequently blocked.

AI Crawler Company Purpose Blocked Blocked % Allowed Allowed %
CCBot Common Crawl Training 187 9.9% 16 0.8%
ClaudeBot Anthropic Training 176 9.3% 26 1.4%
GPTBot OpenAI Training 174 9.2% 38 2.0%
Bytespider ByteDance Training 173 9.1% 4 0.2%
meta-externalagent Meta Training 156 8.2% 5 0.3%
Applebot-Extended Apple Training 153 8.1% 14 0.7%
PerplexityBot Perplexity Search 134 7.1% 40 2.1%
Diffbot Diffbot Extraction 129 6.8% 2 0.1%
cohere-ai Cohere Training 118 6.2% 5 0.3%
Google-Extended Google Training 113 6.0% 32 1.7%
Amazonbot Amazon Training 110 5.8% 11 0.6%
OAI-SearchBot OpenAI Search 106 5.6% 39 2.1%
Claude-User Anthropic Retrieval 84 4.4% 17 0.9%
ChatGPT-User OpenAI Retrieval 79 4.2% 41 2.2%
FacebookBot Meta Preview 79 4.2% 3 0.2%
View data table
AI crawler access policy distribution in robots.txt
Policy Domains Percentage
Blocks All AI 17 0.9%
Blocks Selectively 239 12.6%
Rate-Limits AI 6 0.3%
Explicitly Allows 38 2.0%
No AI Policy 1,596 84.2%

File Quality Distribution

Among the ADF files that were found, how many meet the full specification versus providing only minimal content or containing errors. Quality is assessed using per-file structural checks: required fields must pass for a file to be considered valid; recommended fields distinguish "complete" from "minimal" implementations.

52.4%
43.3%
Complete 52.4% Minimal 4.3% Invalid 43.3%

AI Readiness Tiers

Each domain receives a readiness tier from 0 (Unaware) to 5 (AI-Optimised) based on three inputs: valid ADF file count, AI crawler policy in robots.txt, and Schema.org presence on the homepage. The tier model is deterministic with no opaque weights: the full calculation logic is published.

View data table
AI readiness tier distribution (average score: 2.2 / 5.0)
Tier Domains Percentage
Tier 5: AI-Optimised 0 0.0%
Tier 4: AI-Ready 61 3.2%
Tier 3: Partially Ready 355 18.7%
Tier 2: Passive 1,417 74.7%
Tier 1: Actively Blocking 17 0.9%
Tier 0: Unaware 46 2.4%

ADF vs Other Web Standards

Comparing AI Discovery File adoption against established web standards. This contextualises where ADF adoption sits relative to conventions like robots.txt (RFC 9309), ads.txt (IAB Tech Lab), security.txt (RFC 9116), and humans.txt, all of which also require placing files at the domain root.

View data table
AI Discovery File adoption compared with established web standards
Standard Adoption
robots.txt 49.9%
ads.txt 16.1%
Schema.org 27.6%
security.txt 14.5%
humans.txt 2.3%
Any ADF file 10.2%

Notable Adopters

The top 20 domains by AI readiness tier, showing which high-profile websites are leading ADF adoption. Readiness tiers are calculated using the combinatorial scoring model.

Domain Rank Category Files Found Files Valid Readiness
adobe.com 68 Global Top 1,000 1 1 AI-Ready
asus.com 710 Global Top 1,000 1 1 AI-Ready
bmmagazine.co.uk 739 UK Top 1,000 1 1 AI-Ready
branch.io 429 Global Top 1,000 1 1 AI-Ready
calendly.com 428 Global Top 1,000 2 1 AI-Ready
checkpoint.com 327 Global Top 1,000 1 1 AI-Ready
classlink.com 854 Global Top 1,000 1 1 AI-Ready
cloudflare.com 7 Global Top 1,000 1 1 AI-Ready
cloudinary.com 725 Global Top 1,000 1 1 AI-Ready
datadoghq.com 788 Global Top 1,000 1 1 AI-Ready
dell.com 368 Global Top 1,000 1 1 AI-Ready
dynatrace.com 546 Global Top 1,000 1 1 AI-Ready
dyson.co.uk 462 UK Top 1,000 1 1 AI-Ready
energysavingtrust.org.uk 489 UK Top 1,000 1 1 AI-Ready
exeter.ac.uk 161 UK Top 1,000 1 1 AI-Ready
foxnews.com 450 Global Top 1,000 1 1 AI-Ready
frontiersin.org 917 Global Top 1,000 1 1 AI-Ready
greenpeace.org.uk 724 UK Top 1,000 1 1 AI-Ready
gumgum.com 722 Global Top 1,000 1 1 AI-Ready
hcaptcha.com 474 Global Top 1,000 1 1 AI-Ready

Download the Data

Raw datasets from this quarter's crawl, licensed under CC BY 4.0. Use them for your own research, analysis, or reporting. When citing, please reference the quarter (e.g., "Q4 2026") and link to the methodology.

Methodology

How We Collect This Data

Our crawler checks the top 1,000 global and top 1,000 UK domains (deduplicated to ~1,995) for all 10 AI Discovery Files, validates each against the specification, analyses robots.txt AI crawler policies across 15 known agents, and scores each domain's overall AI readiness using a deterministic tier model. The full methodology, including validation rules, soft 404 detection, redirect classification, and scoring logic, is published for transparency.

Full methodology