Research · First edition
We measured 176 online businesses that already win on Google, and asked a narrower question: can an AI assistant read them at all?
of sites returned nothing at all to a research crawler that identified itself by name and published a contact URL. Every one of them ranks on the first page of Google for a commercial search.
| Market | Attempted | Returned nothing | Share |
|---|---|---|---|
| United Kingdom | 103 | 24 | 23% |
| United States | 73 | 25 | 34% |
| Both | 176 | 49 | 28% |
These businesses let Google in and returned nothing to us. Whatever the mechanism, the practical effect is the same: a crawler that announces itself gets no content, and content that cannot be fetched cannot be read, summarised or cited.
Update, 17 August 2026. The second edition is published, and it corrects this number. We re-requested all 49 sites twice, as an identified crawler and as a browser. Only 5 refuse the crawler while serving the browser, and 39% now respond normally to the same crawler that got nothing on 7 August. Read the second edition →
We then turned the same question on the people selling the fix. Of 81 agencies that sell AI visibility, answer engine optimisation or generative engine optimisation, 61 appear in none of the buying questions their own prospective clients would ask — and not one had its own website read as a source. Read the agency edition →
Of the 127 sites that did return data, these are the checks that came back missing or incomplete. Margins are 95% confidence intervals.
| Missing or incomplete | Share | Margin |
|---|---|---|
| Structured data detail | 43.3% | ± 8.6 pp |
| MCP server | 42.5% | ± 8.6 pp |
| Business identity (Organization / WebSite) | 41.7% | ± 8.6 pp |
llms.txt | 40.2% | ± 8.5 pp |
| Product or service schema | 38.6% | ± 8.5 pp |
| Open Graph | 32.3% | ± 8.1 pp |
| Policies visible to a crawler | 26.8% | ± 7.7 pp |
This is the part that argues against our own product category, so we are putting it in the report rather than in a footnote.
For 64 of the stores we also measured appearance: how often the business showed up when an AI assistant was asked real buying questions in its own category. This is the extension promised in the first edition of this page, which reported 25.
| Technical score | n | Mean appearance | Appeared at all |
|---|---|---|---|
| 90 and above | 51 | 10.8% | 12 of 51 |
| Below 90 | 13 | 14.1% | 4 of 13 |
The larger sample reversed the direction. With 25 stores the better-scoring group appeared more often, and we wrote here that this was the direction anyone would expect. With 64 it is the other way round: the stores scoring below 90 appear more often than the stores scoring above it. We are leaving that sentence in the record rather than quietly deleting it.
Neither gap is large enough to carry a claim on its own. What is harder to explain away are the individual cases:
We expected the United States to be well ahead. It is not, and the difference is not measurable at this sample size.
| Market | n | Mean score | Std. deviation |
|---|---|---|---|
| United Kingdom | 51 | 87.4 | 17.0 |
| United States | 41 | 88.2 | 15.7 |
The difference is 0.8 points, 95% CI [−5.9, +7.5]. Zero sits comfortably inside that interval. Treat the two markets as the same number.
Shopify sites score 95.0 and everything else 58.9, but roughly 17 of those 36 points are protocol files Shopify publishes automatically. Reporting "95 versus 59" without that decomposition would be misleading, so we are not reporting it as a headline. 71 of the 97 stores in this sample are Shopify, which means every aggregate number here carries that platform's built-in advantage.
Frame. Organic results of commercial searches, 10 categories across 2 markets, with marketplaces and large chains excluded so that the sample is the businesses an assistant would have to choose between.
Collection. Each site was fetched by a crawler
identifying itself as GoppaResearchBot, pointing at
trygoppa.com/bot, which explains the study and gives a
removal contact. 176 attempts, 127 with enough data to score: 97 stores
and 30 service businesses.
Scoring. 17 checks across retrieval, readability, identity and discovery. Checks that do not apply to a site type are excluded rather than failed. Scoring a plumber on GTINs would produce a number that means nothing.
Reproducibility. Every percentage on this page except the non-response rate can be recomputed from the CSV below. The non-response rate cannot: sites that returned nothing have no row in a dataset of measurements. That count comes from the collection log, and we say so rather than quietly omitting it.
The full per-site dataset, with domains removed. One row per scored site, one column per check. Free to use, including commercially, with attribution.
Domains are withheld on purpose. The point of the study is the shape of the market, not a list of businesses to embarrass, and several of the sites here are invisible to AI assistants without anyone there knowing it. If you are a researcher who needs the domains for replication, write to us.
The aggregate hides what is happening inside each category. These go deeper on one at a time, including which sources the assistant read instead of the retailers.
Trustpilot was the most-cited source in both. Two categories is not a pattern, and we will say whether it holds when the rest are measured.
Done since the first edition: the appearance measurement was extended from 25 stores to 64, and the table above is that larger sample. It reversed the direction of the score comparison, which is written up where it happened rather than here.
Still open: HTTP status codes logged, so non-response can be attributed instead of guessed; and at least one non-English market, because everything above stops at the Channel.