WordPress
WordPress robots.txt blocking pages: fix it
Work out whether WordPress, a plugin, or a physical robots.txt owns the live rule, then remove the block from the copy that is served.
WordPress normally generates a virtual robots.txt, but a physical root file, plugin, theme, or "robots_txt" filter can replace its rules. Identify which copy serves the unwanted Disallow, remove it there, clear any cache, and test the live file again. Do not use robots.txt noindex as a substitute.
Why WordPress does this
WordPress core generates a virtual robots.txt when no physical robots.txt exists. Its default output disallows wp-admin and allows wp-admin/admin-ajax.php; it does not block normal posts or pages. Plugins and themes can change that virtual output through the robots_txt filter, while a physical robots.txt in the site root can take precedence. A Disallow affects crawling, not guaranteed removal from search. Google can still list the URL without crawling its content.
Check it right now
Before changing anything, confirm what a crawler actually sees. The check is free, takes one URL and needs no account.
How to fix it
- Run the robots.txt tester on the affected URL and copy the exact matching rule from the live response.
- Check the WordPress installation root for a physical file named robots.txt. If it exists, edit that file and remove only the matching "Disallow".
- If no physical file exists, WordPress is serving a virtual response. Check the active SEO plugin and theme code for the matching rule or a robots_txt filter, and remove it at its source.
- If the virtual owner cannot be identified, create a physical robots.txt in the site root using the safe lines from the current response but without the bad rule. WordPress support documents this as the way to override the virtual file.
- Clear any page or edge cache, then run the tester again against the public URL.
- Do not put noindex in robots.txt. Google does not support it there, and a crawler blocked by robots.txt cannot see a noindex on the page.
Why it happens again
WordPress can have two apparent owners for one public response. Editing a plugin-generated virtual file does nothing while a physical root file is served, and editing the root file does nothing if a server or plugin intercepts the request first. Test the public URL after every change.
stillindexed re-checks the URLs you give it every 30 minutes on Starter and alerts when a directive changes, at most 30 minutes after it does. It is a monitor rather than a crawler: it watches a list you choose and tells you when one of seven things changes. Card first, no trial, and a 30 day refund.
Catching it next time
Fixing it once is the easy half. The setting that caused this can be changed again by anyone with access, and the page will keep returning 200 while it happens.
Other ways WordPress loses pages
- my WordPress staging site is indexed by Google
- Every page on the site vanished from Google, and the last thing anyone touched was the WordPress dashboard.
- A WordPress page's canonical tag points at the wrong URL, usually because two SEO plugins, a theme or a code filter disagree about who writes it.
- my WordPress migration lost pages from search
- The http, www or non-www version of a WordPress site takes an extra hop before reaching the final URL, because the host and WordPress each redirect once.
robots.txt blocking pages that should be crawled, on other platforms
Sources
Every claim about WordPress above is from their own documentation, read on 2026-08-30. Platforms change their settings; if one of these is out of date, their page wins and we would like to know.
- https://developer.wordpress.org/reference/functions/do_robots/
- https://developer.wordpress.org/reference/hooks/robots_txt/
- https://developer.wordpress.org/reference/functions/do_robots/
- https://developers.google.com/search/docs/crawling-indexing/robots/intro
- https://developers.google.com/search/docs/crawling-indexing/block-indexing