Research · 2026-10-04
How 85 Major Sites Treat AI Crawlers
We checked robots.txt, llms.txt and homepage structured data on 85 well-known SaaS, media, ecommerce and AI sites. Here is what stood out.
Key findings
36 of 85 sites (42%) publish an llms.txt file. Adoption is concentrated in software companies: 32 of 49 SaaS sites have one, against 0 of 21 media sites.
13 of 72 readable robots.txt files (18%) block GPTBot, OpenAI's training crawler. Media sites account for most of them (8 of 12).
7 of 72 (10%) block at least one AI search crawler (OAI-SearchBot, PerplexityBot or Claude-SearchBot). Blocking those is the choice that removes a site from that engine's answers.
40 of 62 homepages (65%) declare Organization-type JSON-LD, and only 5 (8%) have FAQPage markup on the homepage.
By category
| Category | Sites | Publish llms.txt | Block GPTBot | Block an AI search crawler | Organization schema |
|---|---|---|---|---|---|
| SaaS | 49 | 32 (65%) | 3 of 45 | 0 of 45 | 32 of 44 |
| Media | 21 | 0 (0%) | 8 of 12 | 5 of 12 | 4 of 8 |
| Ecommerce | 10 | 3 (30%) | 2 of 10 | 2 of 10 | 3 of 7 |
| AI companies | 5 | 1 (20%) | 0 of 5 | 0 of 5 | 1 of 3 |
Sites that block an AI search crawler
cnn.com, forbes.com, washingtonpost.com, bloomberg.com, quora.com, amazon.com, ebay.com.
Whether that is a deliberate licensing decision or a leftover rule, the effect is the same: the blocked engine cannot fetch those pages for its answers. Check your own file with the AI crawler checker.
What to take from this
Publishers often block on purpose. A business that wants to be recommended should not. If your robots.txt came from a template, a CDN default or an old security rule, confirm that it lets AI search crawlers in, then build a clean one with the robots.txt generator. Publish an llms.txt and add Organization schema to your homepage.
Method and limits
How was the data collected?
On 2026-10-04 we fetched each site's /robots.txt, /llms.txt and homepage over plain HTTPS from one server, and parsed them with the same robots.txt rules AEOCheck's scanner uses. The script is in the AEOCheck repository (scripts/ai-readiness-study.mjs) and the raw results are in data/ai-readiness-study.json.
Is this a representative sample?
No. The list is 85 well-known sites we picked by hand across four categories. It shows how prominent sites behave, not the whole web. Sites whose robots.txt or homepage our request could not read (bot protection, redirects, errors) are excluded from the matching percentages, which is why the denominators differ.
Does blocking an AI crawler mean a site is not cited?
Not necessarily. Robots.txt is one signal. Some systems also use licensed data or other crawlers, and robots.txt is advisory. Blocking an AI search crawler does remove the site from that engine's own crawl, which is the relevant risk for visibility.
Does having llms.txt improve AI visibility?
We make no such claim. The data shows who publishes the file, not what it earns them. It is a low-cost way to give AI systems a summary of your site, and the free llms.txt generator builds one in a minute.
Want to cite this? Link to this page and name AEOCheck and the date above.