Research · First edition
We measured 176 online businesses that already win on Google, and asked a narrower question: can an AI assistant read them at all?
of sites returned nothing at all to a research crawler that identified itself by name and published a contact URL. Every one of them ranks on the first page of Google for a commercial search.
| Market | Attempted | Returned nothing | Share |
|---|---|---|---|
| United Kingdom | 103 | 24 | 23% |
| United States | 73 | 25 | 34% |
| Both | 176 | 49 | 28% |
These businesses let Google in and returned nothing to us. Whatever the mechanism, the practical effect is the same: a crawler that announces itself gets no content, and content that cannot be fetched cannot be read, summarised or cited.
Of the 127 sites that did return data, these are the checks that came back missing or incomplete. Margins are 95% confidence intervals.
| Missing or incomplete | Share | Margin |
|---|---|---|
| Structured data detail | 43.3% | ± 8.6 pp |
| MCP server | 42.5% | ± 8.6 pp |
| Business identity (Organization / WebSite) | 41.7% | ± 8.6 pp |
llms.txt | 40.2% | ± 8.5 pp |
| Product or service schema | 38.6% | ± 8.5 pp |
| Open Graph | 32.3% | ± 8.1 pp |
| Policies visible to a crawler | 26.8% | ± 7.7 pp |
This is the part that argues against our own product category, so we are putting it in the report rather than in a footnote.
For 25 of the stores we also measured appearance: how often the business showed up when an AI assistant was asked real buying questions in its own category.
| Technical score | n | Mean appearance | Appeared at all |
|---|---|---|---|
| 90 and above | 15 | 6.7% | 3 of 15 |
| Below 90 | 10 | 3.4% | 2 of 10 |
The direction is the one you would expect, and the sample is far too small to carry it. What is harder to explain away are the individual cases:
We expected the United States to be well ahead. It is not, and the difference is not measurable at this sample size.
| Market | n | Mean score | Std. deviation |
|---|---|---|---|
| United Kingdom | 51 | 87.4 | 17.0 |
| United States | 41 | 88.2 | 15.7 |
The difference is 0.8 points, 95% CI [−5.9, +7.5]. Zero sits comfortably inside that interval. Treat the two markets as the same number.
Shopify sites score 95.0 and everything else 58.9 — but roughly 17 of those 36 points are protocol files Shopify publishes automatically. Reporting "95 versus 59" without that decomposition would be misleading, so we are not reporting it as a headline. 71 of the 97 stores in this sample are Shopify, which means every aggregate number here carries that platform's built-in advantage.
Frame. Organic results of commercial searches, 10 categories across 2 markets, with marketplaces and large chains excluded so that the sample is the businesses an assistant would have to choose between.
Collection. Each site was fetched by a crawler
identifying itself as GoppaResearchBot, pointing at
trygoppa.com/bot, which explains the study and gives a
removal contact. 176 attempts, 127 with enough data to score: 97 stores
and 30 service businesses.
Scoring. 17 checks across retrieval, readability, identity and discovery. Checks that do not apply to a site type are excluded rather than failed — scoring a plumber on GTINs would produce a number that means nothing.
Reproducibility. Every percentage on this page except the non-response rate can be recomputed from the CSV below. The non-response rate cannot: sites that returned nothing have no row in a dataset of measurements. That count comes from the collection log, and we say so rather than quietly omitting it.
The full per-site dataset, with domains removed. One row per scored site, one column per check. Free to use, including commercially, with attribution.
Domains are withheld on purpose. The point of the study is the shape of the market, not a list of businesses to embarrass — and several of the sites here are invisible to AI assistants without anyone there knowing it. If you are a researcher who needs the domains for replication, write to us.
The aggregate hides what is happening inside each category. These go deeper on one at a time, including which sources the assistant read instead of the retailers.
Trustpilot was the most-cited source in both. Two categories is not a pattern, and we will say whether it holds when the rest are measured.
Three things will change: HTTP status codes logged, so non-response can be attributed instead of guessed; the appearance measurement extended from 25 stores to the full sample; and at least one non-English market, because everything above stops at the Channel.