goppa

Research · Second edition

Why 28% Returned Nothing

The first edition said we could not name a cause, and promised this. We went back to the 49 sites that returned nothing and requested each one twice — changing only who we said we were.

Re-measured 17 August 2026 49 sites · 98 requests United Kingdom & United States Dataset published in full

The finding

5 of 49

refuse a crawler that identifies itself while serving a browser the same second. That is 10% of the sites we could not read — not 28% of the internet, and not what "AI crawlers are being blocked" is usually taken to mean.

The other 44 split into three groups, and only one of them is about us at all.

What the two requests showedSitesShare
Refused the crawler, served the browser
The block is on identity
510%
Refused both
Not about who we said we were
2041%
Served both
Responding normally today
1939%
No response either way
Down, or unreachable from here
510%

The part that surprised us most

39% of the sites that returned nothing on 7 August returned a normal page on 17 August — to the same crawler, with the same name and the same contact URL. Nothing about our request changed.

That means a single measurement of crawler access is not a property of a website. It is a snapshot. Ours was, too. Any figure published from one pass — including the 28% in our own first edition — describes a moment, and reporting it as a characteristic of the site overstates what was measured.

This corrects our own first edition. We published 28% and said we could not name the cause. We were right not to guess: had we assumed the cause, we would have attributed to deliberate AI blocking a number that is mostly something else. The honest version of that headline is narrower and less dramatic — 10% of the unreadable sites refuse an identified crawler while serving a browser.

What we still cannot say

Four limits, and they matter more than the numbers above.

LimitWhy it matters
Both requests left the same address"Refused both" may be datacenter-IP reputation rather than the site refusing everyone. A browser on a home connection could well get through. We cannot separate the two from here
This is not a test of GPTBotWe measured what happens to our crawler. Many bot-protection products allow GPTBot and Googlebot by name and refuse everything else. A site that refuses us may well serve ChatGPT's crawler
The sample is the failuresThese 49 are the sites that failed in the first edition, not a fresh random sample. The rates here describe that group and nothing wider
Ten days apartConfiguration changes, rate limits and incidents all live inside that window. We report both dates rather than pick one

Who answered the requests

Among the 25 sites that refused us in either form, one product accounts for the majority of the responses.

Server headerSites
Cloudflare15
CloudFront2
Sucuri2
nginx2
Akamai, Amazon S3, DataDome, unidentified4

We are not saying Cloudflare blocks AI crawlers. Cloudflare is the most widely deployed product in this list, so it appears most often for the same reason it appears most often everywhere. What the header tells you is where the setting lives — and in every one of the five identity refusals, it is a setting the site owner can change.

By market

MarketRe-checkedRefused crawler, served browserResponding today
United Kingdom2437
United States25212
Both49519

With five cases in total, the difference between markets is not a difference. We publish the split because withholding it would be worse, not because it means anything.

Method

Two HTTP GET requests per site, seconds apart, from the same address. The only difference between them is the User-Agent header:

We record both status codes and the Server response header, and we assign a conclusion only where the pair of statuses supports one. No conclusion is inferred from a single request. The full table, one row per site, is published below.

If you run one of these sites and want to check it yourself: request your own homepage with a bot-like User-Agent and then with a browser one. If the first is refused and the second is not, the setting is in your WAF, and allowing named AI crawlers is usually one rule.

Download the dataset (CSV, 49 rows) Read the first edition

How to cite: Goppa Research (2026). Why 28% Returned Nothing: crawler access re-checked, second edition. Published 17 August 2026. https://trygoppa.com/study/2 — CC BY 4.0.

Domains are not published. The dataset carries market, category, both status codes, the server header and the conclusion — every field except the identity of the business. We are not going to name a company for a setting it probably did not choose.