Research · Second edition
The first edition said we could not name a cause, and promised this. We went back to the 49 sites that returned nothing and requested each one twice — changing only who we said we were.
refuse a crawler that identifies itself while serving a browser the same second. That is 10% of the sites we could not read — not 28% of the internet, and not what "AI crawlers are being blocked" is usually taken to mean.
The other 44 split into three groups, and only one of them is about us at all.
| What the two requests showed | Sites | Share |
|---|---|---|
| Refused the crawler, served the browser The block is on identity | 5 | 10% |
| Refused both Not about who we said we were | 20 | 41% |
| Served both Responding normally today | 19 | 39% |
| No response either way Down, or unreachable from here | 5 | 10% |
39% of the sites that returned nothing on 7 August returned a normal page on 17 August — to the same crawler, with the same name and the same contact URL. Nothing about our request changed.
That means a single measurement of crawler access is not a property of a website. It is a snapshot. Ours was, too. Any figure published from one pass — including the 28% in our own first edition — describes a moment, and reporting it as a characteristic of the site overstates what was measured.
Four limits, and they matter more than the numbers above.
| Limit | Why it matters |
|---|---|
| Both requests left the same address | "Refused both" may be datacenter-IP reputation rather than the site refusing everyone. A browser on a home connection could well get through. We cannot separate the two from here |
| This is not a test of GPTBot | We measured what happens to our crawler. Many bot-protection products allow GPTBot and Googlebot by name and refuse everything else. A site that refuses us may well serve ChatGPT's crawler |
| The sample is the failures | These 49 are the sites that failed in the first edition, not a fresh random sample. The rates here describe that group and nothing wider |
| Ten days apart | Configuration changes, rate limits and incidents all live inside that window. We report both dates rather than pick one |
Among the 25 sites that refused us in either form, one product accounts for the majority of the responses.
| Server header | Sites |
|---|---|
| Cloudflare | 15 |
| CloudFront | 2 |
| Sucuri | 2 |
| nginx | 2 |
| Akamai, Amazon S3, DataDome, unidentified | 4 |
We are not saying Cloudflare blocks AI crawlers. Cloudflare is the most widely deployed product in this list, so it appears most often for the same reason it appears most often everywhere. What the header tells you is where the setting lives — and in every one of the five identity refusals, it is a setting the site owner can change.
| Market | Re-checked | Refused crawler, served browser | Responding today |
|---|---|---|---|
| United Kingdom | 24 | 3 | 7 |
| United States | 25 | 2 | 12 |
| Both | 49 | 5 | 19 |
With five cases in total, the difference between markets is not a difference. We publish the split because withholding it would be worse, not because it means anything.
Two HTTP GET requests per site, seconds apart, from the same address.
The only difference between them is the User-Agent header:
We record both status codes and the Server response header,
and we assign a conclusion only where the pair of statuses supports one.
No conclusion is inferred from a single request. The full table, one row
per site, is published below.
If you run one of these sites and want to check it yourself: request your own homepage with a bot-like User-Agent and then with a browser one. If the first is refused and the second is not, the setting is in your WAF, and allowing named AI crawlers is usually one rule.
How to cite: Goppa Research (2026). Why 28% Returned Nothing: crawler access re-checked, second edition. Published 17 August 2026. https://trygoppa.com/study/2 — CC BY 4.0.
Domains are not published. The dataset carries market, category, both status codes, the server header and the conclusion — every field except the identity of the business. We are not going to name a company for a setting it probably did not choose.