10.2% adoption
AI Discovery File Adoption Research
Measuring how the world's top websites prepare for AI systems. Based on a crawl of 1,995 domains.
Summary
This Q4 2026 report analyses AI Discovery File adoption across 1,896 of the web's most prominent domains. 10.2% of domains have at least one AI Discovery File, 84.2% have no AI-specific crawler policy in their robots.txt, and the average AI readiness tier is 2.2 out of 5.0. Data is collected quarterly using the methodology described in our full methodology documentation.
AI Discovery Files are a set of 10 standardised root-level files, including llms.txt, ai.txt, ai.json, identity.json, and brand.txt, that help AI systems such as ChatGPT, Claude, and Gemini discover, interpret, and correctly represent a website. This research tracks their real-world adoption among the domains most likely to be referenced by AI systems when answering user questions.
Key Findings
The Q4 2026 crawl puts AI Discovery File adoption above 10% for the first time: 194 of 1,896 successfully crawled domains (10.2%) now publish at least one file, up from 164 (9.4%) in Q3. llms.txt still carries almost all of it, but the files it carries are getting better, and the number of AI-Ready sites rose by more than a third. This quarter is also the cleanest comparison so far. Only 99 domains errored, against 251 in July, so the sample is 152 domains larger and much closer to the full list. That cuts both ways: the percentages are more trustworthy, and some of the new adopters may simply be sites we could not reach last time rather than sites that changed.
- Adoption passes 10%, and the better crawl makes the rise easier to trust Domains with any AI Discovery File rose from 164 to 194 (+18%), and the share from 9.4% to 10.2%. The share matters more than the count here, because the crawl reached 1,896 domains this quarter against 1,744 in July (errors fell from 251 to 99). A rise in the percentage on a larger, more complete sample is the opposite of last quarter, when a shrinking sample flattered some numbers. We still cannot tell from aggregate data alone how many of the 30 new adopters were unreachable in July rather than newly published.
- llms.txt grows, and its quality improves faster than its count llms.txt was found on 150 domains, up from 122 (+23%), while valid and complete files rose from 82 to 110 (+34%). Invalid llms.txt files stayed exactly flat at 40, so every additional file this quarter was a good one: the valid share rose from 67% to 73%. Across all file types the count of complete files went from 85 to 110. Publishers are not only adopting llms.txt, they are getting the format right more often.
- AI-Ready sites up by more than a third, with Calendly, Check Point and Exeter reaching it for the first time Tier 4 (AI-Ready) domains rose from 44 to 61 (+39%), now 3.2% of the sample against 2.5% in Q3. Calendly, Check Point, Branch, hCaptcha and the University of Exeter reach AI-Ready for the first time in our data, and Dell and Greenpeace UK, both AI-Ready in Q2, are back after missing the July list. Movement runs both ways: Bunkbedsstore and Groupon UK still publish a valid llms.txt but fell to Tier 3, and Bluehost now returns 403 to our crawler while serving the same file to GPTBot and ClaudeBot (see the methodology note on user-agent gating). Tier 5 (AI-Optimised) is still empty after four quarters, and the average readiness score held at 2.2 out of 5.0.
- Blanket AI blocking falls, explicit permission rises again Sites blocking every AI crawler fell from 23 to 17 (1.3% to 0.9%), the first real drop in four quarters, while sites explicitly allowing AI crawlers rose from 31 to 38 (2.0%). Selective blocking edged up from 217 to 239 domains (12.6%), so the trend is towards deliberate, per-crawler decisions rather than all-or-nothing. The large majority is unchanged: 84.2% of sites still have no AI policy at all, the same share as in July.
- Other file types appear for the first time, but none of them is valid yet identity.json, brand.txt and developer-ai.txt were each found on a domain for the first time, and ai.txt and robots-ai.txt each on two. None of those new files passed validation, so they are attempts rather than adoption. llms.html moved the wrong way on quality: found on 49 domains (from 46), but valid files fell from 9 to 7 and complete files from 3 to 0. Breadth is starting, but only llms.txt is being implemented to a standard a machine can rely on.
- Still zero of 194 adopters reference a formal specification For the fourth quarter running, not one domain publishing an AI Discovery File references a formal specification. For scale, AI Discovery Files are now on 10.2% of crawled sites against 14.5% for security.txt and 2.3% for humans.txt, so this is no longer a fringe practice, yet every one of these files is still written without a shared schema or version to validate against. Adoption keeps outrunning standardisation, and the gap widens as adoption grows.
AI Visibility Research, October 2026
Changes from Q3 2026
Quarter-over-quarter changes in key metrics between Q3 2026 and Q4 2026.
| File | Q3 2026 | Q4 2026 | Change |
|---|---|---|---|
| llms.txt | 7.0% | 7.9% | +0.9 |
| llms.html | 2.6% | 2.6% | 0.0 |
| ai.txt | 0.1% | 0.1% | 0.0 |
| ai.json | 0.1% | 0.1% | 0.0 |
| identity.json | 0.0% | 0.1% | +0.1 |
| brand.txt | 0.0% | 0.1% | +0.1 |
| faq-ai.txt | 0.0% | 0.0% | 0.0 |
| developer-ai.txt | 0.0% | 0.1% | +0.1 |
| robots-ai.txt | 0.1% | 0.1% | 0.0 |
New to This Quarter's Top 20
- branch.io
- calendly.com
- checkpoint.com
- dell.com
- exeter.ac.uk
- greenpeace.org.uk
- hcaptcha.com
Outside This Quarter's Top 20
- bluehost.com
- bunkbedsstore.uk
- groupon.co.uk
- hostgator.com.br
- kingsfund.org.uk
- klaviyo.com
- life360.com
Ranked by readiness tier, then valid files. Most AI-Ready sites tie and ties fall alphabetically, so a site here may still be AI-Ready, and a site in the other list may be returning rather than new.
ADF Adoption by File Type
How many of the top websites have each AI Discovery File: and how many of those files pass structural validation. Files are checked at their canonical root-level URL (e.g., example.com/llms.txt) and validated against the ADF specification.
View data table
| File | Found | Valid | Complete |
|---|---|---|---|
| llms.txt | 150 | 110 | 110 |
| llms.html | 49 | 7 | 0 |
| ai.txt | 2 | 0 | 0 |
| ai.json | 2 | 0 | 0 |
| identity.json | 1 | 0 | 0 |
| brand.txt | 1 | 0 | 0 |
| faq-ai.txt | 0 | 0 | 0 |
| developer-ai.txt | 1 | 0 | 0 |
| robots-ai.txt | 2 | 1 | 0 |
AI Crawler Access Policies
How websites use robots.txt to manage access for 15 known AI user agents: from OpenAI's GPTBot to Anthropic's ClaudeBot. Each domain is classified into one of five access policies based on its aggregate behaviour across all agents. The per-agent table below shows which AI crawlers are most frequently blocked.
| AI Crawler | Company | Purpose | Blocked | Blocked % | Allowed | Allowed % |
|---|---|---|---|---|---|---|
| CCBot | Common Crawl | Training | 187 | 9.9% | 16 | 0.8% |
| ClaudeBot | Anthropic | Training | 176 | 9.3% | 26 | 1.4% |
| GPTBot | OpenAI | Training | 174 | 9.2% | 38 | 2.0% |
| Bytespider | ByteDance | Training | 173 | 9.1% | 4 | 0.2% |
| meta-externalagent | Meta | Training | 156 | 8.2% | 5 | 0.3% |
| Applebot-Extended | Apple | Training | 153 | 8.1% | 14 | 0.7% |
| PerplexityBot | Perplexity | Search | 134 | 7.1% | 40 | 2.1% |
| Diffbot | Diffbot | Extraction | 129 | 6.8% | 2 | 0.1% |
| cohere-ai | Cohere | Training | 118 | 6.2% | 5 | 0.3% |
| Google-Extended | Training | 113 | 6.0% | 32 | 1.7% | |
| Amazonbot | Amazon | Training | 110 | 5.8% | 11 | 0.6% |
| OAI-SearchBot | OpenAI | Search | 106 | 5.6% | 39 | 2.1% |
| Claude-User | Anthropic | Retrieval | 84 | 4.4% | 17 | 0.9% |
| ChatGPT-User | OpenAI | Retrieval | 79 | 4.2% | 41 | 2.2% |
| FacebookBot | Meta | Preview | 79 | 4.2% | 3 | 0.2% |
View data table
| Policy | Domains | Percentage |
|---|---|---|
| Blocks All AI | 17 | 0.9% |
| Blocks Selectively | 239 | 12.6% |
| Rate-Limits AI | 6 | 0.3% |
| Explicitly Allows | 38 | 2.0% |
| No AI Policy | 1,596 | 84.2% |
File Quality Distribution
Among the ADF files that were found, how many meet the full specification versus providing only minimal content or containing errors. Quality is assessed using per-file structural checks: required fields must pass for a file to be considered valid; recommended fields distinguish "complete" from "minimal" implementations.
AI Readiness Tiers
Each domain receives a readiness tier from 0 (Unaware) to 5 (AI-Optimised) based on three inputs: valid ADF file count, AI crawler policy in robots.txt, and Schema.org presence on the homepage. The tier model is deterministic with no opaque weights: the full calculation logic is published.
View data table
| Tier | Domains | Percentage |
|---|---|---|
| Tier 5: AI-Optimised | 0 | 0.0% |
| Tier 4: AI-Ready | 61 | 3.2% |
| Tier 3: Partially Ready | 355 | 18.7% |
| Tier 2: Passive | 1,417 | 74.7% |
| Tier 1: Actively Blocking | 17 | 0.9% |
| Tier 0: Unaware | 46 | 2.4% |
ADF vs Other Web Standards
Comparing AI Discovery File adoption against established web standards. This contextualises where ADF adoption sits relative to conventions like robots.txt (RFC 9309), ads.txt (IAB Tech Lab), security.txt (RFC 9116), and humans.txt, all of which also require placing files at the domain root.
View data table
| Standard | Adoption |
|---|---|
| robots.txt | 49.9% |
| ads.txt | 16.1% |
| Schema.org | 27.6% |
| security.txt | 14.5% |
| humans.txt | 2.3% |
| Any ADF file | 10.2% |
Notable Adopters
The top 20 domains by AI readiness tier, showing which high-profile websites are leading ADF adoption. Readiness tiers are calculated using the combinatorial scoring model.
| Domain | Rank | Category | Files Found | Files Valid | Readiness |
|---|---|---|---|---|---|
| adobe.com | 68 | Global Top 1,000 | 1 | 1 | AI-Ready |
| asus.com | 710 | Global Top 1,000 | 1 | 1 | AI-Ready |
| bmmagazine.co.uk | 739 | UK Top 1,000 | 1 | 1 | AI-Ready |
| branch.io | 429 | Global Top 1,000 | 1 | 1 | AI-Ready |
| calendly.com | 428 | Global Top 1,000 | 2 | 1 | AI-Ready |
| checkpoint.com | 327 | Global Top 1,000 | 1 | 1 | AI-Ready |
| classlink.com | 854 | Global Top 1,000 | 1 | 1 | AI-Ready |
| cloudflare.com | 7 | Global Top 1,000 | 1 | 1 | AI-Ready |
| cloudinary.com | 725 | Global Top 1,000 | 1 | 1 | AI-Ready |
| datadoghq.com | 788 | Global Top 1,000 | 1 | 1 | AI-Ready |
| dell.com | 368 | Global Top 1,000 | 1 | 1 | AI-Ready |
| dynatrace.com | 546 | Global Top 1,000 | 1 | 1 | AI-Ready |
| dyson.co.uk | 462 | UK Top 1,000 | 1 | 1 | AI-Ready |
| energysavingtrust.org.uk | 489 | UK Top 1,000 | 1 | 1 | AI-Ready |
| exeter.ac.uk | 161 | UK Top 1,000 | 1 | 1 | AI-Ready |
| foxnews.com | 450 | Global Top 1,000 | 1 | 1 | AI-Ready |
| frontiersin.org | 917 | Global Top 1,000 | 1 | 1 | AI-Ready |
| greenpeace.org.uk | 724 | UK Top 1,000 | 1 | 1 | AI-Ready |
| gumgum.com | 722 | Global Top 1,000 | 1 | 1 | AI-Ready |
| hcaptcha.com | 474 | Global Top 1,000 | 1 | 1 | AI-Ready |
Download the Data
Raw datasets from this quarter's crawl, licensed under CC BY 4.0. Use them for your own research, analysis, or reporting. When citing, please reference the quarter (e.g., "Q4 2026") and link to the methodology.
Methodology
How We Collect This Data
Our crawler checks the top 1,000 global and top 1,000 UK domains (deduplicated to ~1,995) for all 10 AI Discovery Files, validates each against the specification, analyses robots.txt AI crawler policies across 15 known agents, and scores each domain's overall AI readiness using a deterministic tier model. The full methodology, including validation rules, soft 404 detection, redirect classification, and scoring logic, is published for transparency.