On the night of 2 September 2026 we fetched the homepage, robots.txt and first declared sitemap of the 10000 highest-ranked domains on the Tranco list, once each, with the same parsers stillindexed uses to watch customer sites. 7740 domains answered. 6224 returned a 200 HTML homepage.
This page is what the engine saw, what we did to separate real problems from artefacts of a single fetch, and how to reproduce every line of it. The CSV of all 10000 records is linked at the bottom.
1. Homepages that tell Google to stay out
Raw flags, before any judgement:
- 87 of 6224 200 HTML homepages (1.4%) carried
<meta name="robots">withnoindex. - 10 of 6224 200 HTML homepages (0.2%) sent
noindexin anX-Robots-Tagresponse header. - 143 of 6209 parseable robots.txt files (2.3%) disallowed Googlebot from
/or from the homepage path. The robots file belongs to the host the homepage fetch ended on, which is not always the Tranco domain.
A raw flag is not a finding. We re-fetched every one of the 225 flagged domains a second time the same night and sorted them by hand into four classes:
| Class | Count | What it means |
|---|---|---|
| By design | 122 | Login and SSO portals, app shells, ad and CDN hosts, parked domains, and sites that disallow / while allowing the homepage path. Reddit, Naver and Daum publish Disallow: / on purpose and are here. |
| Fetch artefact | 83 | Bot walls that answer 200 with an empty noindex stub, Cloudflare and publisher 403s, geo and consent interstitials, captcha redirects, and one bug in our own parser (below). |
| Apparent accident | 14 | A public content homepage, reproduced on a second fetch, that says noindex or sits behind User-agent: * / Disallow: /. Six are noindex (five meta, one header). Eight are root disallows. |
| Unclear | 6 | Could be either. Not counted. |
The 14 include a national government portal, two US state government sites, a university homepage, a wealth-management brand, a podcast hosting platform, a news homepage, and a messaging-widget vendor. We are not naming them yet, and the reason is now a correction rather than a plan: see the next look.
On the 14 public-content homepages that remained after curation, the directive was visible in the HTML, response headers, or robots file. A scheduled check can surface that change while the operator still has time to inspect it.
2. Sitemaps that are declared and do not serve
Robots files for 3958 domains declared at least one Sitemap: URL. We fetched the first declared URL with GET.
567 of 3958 (14.3%) did not serve a usable sitemap at the URL their own robots.txt advertised. By response:
- 235 returned HTTP 200 with a body that did not parse as a sitemap or sitemap index.
- 175 returned HTTP 403. Some may be bot gates that admit verified Googlebot; this scan does not establish how verified Googlebot was treated.
- 78 returned HTTP 404 or 410.
- 6 returned a 5xx status, 63 returned no response, and 10 returned another status.
Sitemaps are optional and Google finds URLs by links as well. A declared sitemap that returned 404 pointed at nothing during this fetch. The scan cannot establish how long that response lasted. Example, harmless to name and rechecked before publication: Pinterest declared https://www.pinterest.com/v3_sitemaps/qem_klp_vlm_vase_high_www.pinterest.com.xml; that URL returned HTTP 404.
curl -sI -A 'stillindexed-survey/1 (+https://stillindexed.com/bot)' \
'https://www.pinterest.com/v3_sitemaps/qem_klp_vlm_vase_high_www.pinterest.com.xml'
Published sitemap results use GET. Preliminary HEAD-only results are not included.
3. The page changes when the visitor calls itself Googlebot
We obtained a second response for 7016 eligible final homepage URLs using Google’s published Googlebot user-agent string from a non-Google IP. The denominator excludes default-fetch challenge pages and refetches without a comparable response.
675 of 7016 refetched homepages (9.6%) answered differently: a different status, a different final URL after redirects, a different canonical, or a noindex that appeared or disappeared.
This is not “9.6% of sites block Googlebot.” Google verifies its crawler by reverse DNS, and a CDN is entitled to challenge a spoofed user agent from a residential IP while letting the real one through. The number shows how often two HTTP responses differed when our client changed its advertised user agent to Googlebot. It does not establish what Chrome or verified Googlebot received.
What this found in our own parser
Six US newspaper sites (USA Today and five Gannett titles) were flagged as disallowing Googlebot. They do not. Their robots.txt has a User-agent: Googlebot-News group, and our robots parser matched the token Googlebot as a prefix of Googlebot-News. Per RFC 9309 and Google’s documentation the product token has to match exactly. The parser is fixed, with a test, and the six rows are in the fetch-artefact class above with the bug named in the CSV. We are publishing this because it is the argument for the curation step: a raw flag from any monitor, including ours, is where the work starts.
Method
Source. We started from the 2 September 2026 Tranco list, applied the repository’s published domain and DNS filters, and took the first 10,000 survivors.
Fetch. User agent stillindexed-survey/1 (+https://stillindexed.com/bot). Concurrency 24, at most one in-flight request per host, 10 second timeout, no retries, 2 MiB body cap, 10 redirect hops. Entry order https://<domain>/, then https://www.<domain>/, then http://<domain>/. robots.txt from the host the homepage ended on. First declared sitemap, else /sitemap.xml, by GET. TLS facts from the same connection. One further GET with the Googlebot user agent. Median 7589 ms per responding domain, p90 21767 ms.
Parsers. The product’s own: HTML robots meta and X-Robots-Tag directive parsing, canonical from <link> and from the HTTP Link header, robots.txt evaluation per RFC 9309 with Google’s group-precedence rules, sitemap and sitemap-index parsing. A challenge page is recorded when a 403 or 503 carries cf-mitigated: challenge, a cf-chl body, or a “Just a moment” / “Attention Required” title, or when the canonical points at a reCAPTCHA URL.
Limits. HTTP only, no JavaScript. Homepage, robots.txt and one sitemap URL per domain. One IP, one moment in time. Geo and consent walls are recorded as seen. A spoofed Googlebot user agent is not Googlebot.
Prior work. HTTP Archive’s 2025 Web Almanac SEO chapter is the page-level benchmark for robots.txt status codes, robots directives, canonicals, titles and H1s. Bouchaud and Ramaciotti (arXiv:2510.09031, 2025) measured Googlebot and AI-crawler disallows across the CrUX top million. Ahrefs published AI-bot block rates in 2025 and site-audit issue rates in 2023. SEOmator published sitemap discovery on 109,440 domains in July 2026. In the public studies we reviewed, we did not find one that combined manual curation of homepage-level flags with named follow-up.
Reproduce. Download the CSV: one row per domain with selected fetch fields, every published verdict, and the curation class for the 225 flagged rows. The scanner is the survey command in the stillindexed repository, which is not public yet. Until 17 September 2026 this sentence linked a seo-guard repository that does not exist and returned 404, reintroduced by the same edit described under the next look. Every figure on this page reproduces from the CSV and the curl commands above without it.
The next look
Correction, 17 September 2026. Until today this section said the 14 were emailed on 3 September and that the follow-up would be published on 10 September. Neither was true. The notices were never sent, and no follow-up was published on the 10th. The sentence had in fact been removed before the study was first published, for exactly that reason, and was reintroduced by mistake in an unrelated edit on 9 September. A product that sells the detection of silent, unnoticed regressions published one on its own research page and did not notice for eight days. It is recorded here rather than quietly deleted, because the alternative is the behaviour this study criticises.
What the 14 look like now. The re-check no longer depends on anybody running it. A scheduled pass re-reads all 14 every week and publishes what it saw, including the failures, at api.stillindexed.com/v1/research/tranco-10k/followup. Each domain is fetched once for the homepage following redirects, once for robots.txt, and once more with a Googlebot user agent to check whether a page that answers us is also answering a crawler. The counts below are that endpoint’s, as of the pass at 12:37 UTC on 17 September 2026.
| Verdict | Count | Detail |
|---|---|---|
| Fixed | 1 | The homepage returns index, follow, no X-Robots-Tag, no root disallow, and the Googlebot-agent fetch is also answered 200. It is ml.com, named because the finding is that it was repaired. |
| Unchanged | 11 | The same noindex, or the same User-agent: * / Disallow: /, as on 2 September. Two of these answer a Googlebot-agent request with 403 while answering us with 200. |
| Changed, still refusing crawlers | 1 | The root disallow is gone from a robots.txt that is now 185 bytes, and the edge answers a Googlebot user agent with 403. Nothing was opened up. |
| Unclear | 1 | The apex redirects to a locale path rather than the homepage, so the page we read is not the page the notice is about. A clean fetch of some other page is not evidence about this one, and it is never counted as fixed. |
A 403 is our request being refused, not evidence about what the site tells Google, and it is never read as a fix. The same rule covers the unclear row: this page previously published three of these 14 as fixed, because a locale redirect and a dropped robots line were both read as good news. That was wrong, and the rules that now prevent it are in the endpoint above.
What happens next, with a date we will keep. Notices went out on 17 September 2026 to the three domains that publish a contact we could verify. The remaining ten publish no address of their own; a guess at webmaster@ is not a notice, and an address belonging to their DNS or CDN provider is not their address. On 1 October 2026 this section will name every domain still blocking, will say which were reachable and which were not, and will record any that replied saying the state is deliberate. If that date passes with nothing published, read it the way you would read any other monitor that went quiet.
If you want to know what your own homepage says to Googlebot right now, the indexability checker runs the same parser, free, no account. If you want to know within fifteen minutes when it changes, that is what stillindexed does.