WordPress Robots.txt: What to Block and What Not To

A good WordPress robots.txt blocks very little. Keep the default rule for the admin area, add your sitemap, and block only junk URLs like internal search results. Never block your CSS, JavaScript, images or any page you want to rank. And remember: robots.txt does not remove a page from Google. A noindex tag does that.

That is the short answer. The rest of this guide explains each rule in plain words, shows what to do about AI crawlers like GPTBot and ClaudeBot, and gives you a safe starter file you can copy. If you just want a file now, our free Robots.txt Generator builds one for you in a minute.

What does a WordPress robots.txt file actually do?

A robots.txt file is a small text file at the root of your site, for example yoursite.com/robots.txt. It tells crawlers (the bots that search engines and AI companies send to read websites) which parts of your site they may visit.

It is a set of polite requests, not a lock. Well behaved bots like Googlebot follow it. Bad bots ignore it. So it is never a way to hide private pages.

It also does not control what shows in search results. Google’s robots.txt guide says this directly: a page blocked in robots.txt can still appear in Google if other sites link to it. It just shows up with no description. To keep a page out of results, you use a noindex tag or a password instead.

What does WordPress put in robots.txt by default?

WordPress does not create a real file on your server. It builds a “virtual” one on the fly when a bot asks for it. On a fresh install it looks like this:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

In plain words: all bots, stay out of the admin folder, but you may use admin-ajax.php. That second line matters, because many themes and plugins load content on the public site through that file.

Recent WordPress versions also add a Sitemap: line when the built in sitemap is on. SEO plugins like Rank Math and Yoast usually swap in their own sitemap address. Once you upload a real robots.txt file to your site root, WordPress stops serving the virtual one, so your file must include everything you need.

One setting surprises people. If you tick “Discourage search engines from indexing this site” under Settings, Reading, modern WordPress adds a noindex tag to your pages. It does not rely on robots.txt for this. So a clean robots.txt does not prove your site is visible. Check that box too.

What to block in your WordPress robots.txt

Block only URLs that waste a crawler’s time and never deserve to rank. For most small business sites, that is a short list:

  • The admin area. /wp-admin/ is already blocked by default. Keep the admin-ajax.php exception.
  • Internal search results. URLs like /?s=shoes can create endless thin pages. Blocking /?s= and /search/ stops bots from crawling them.
  • Endless filter and sort URLs. On WooCommerce shops, filter combinations like ?orderby= or ?filter_color= can multiply into thousands of near copies. Block the parameters you know are junk.
  • Your staging site, completely. A test copy of your site should carry Disallow: / and a password. Just make sure that rule never gets copied to the live site.

What about wp-login.php? Blocking it in robots.txt does no harm, but it adds no security either. Anyone can still open the login page. Real protection comes from strong passwords for business accounts, two step login and limiting login attempts.

What should you never block in robots.txt?

This is where most damage happens. Old WordPress guides told people to block /wp-includes/ and /wp-content/. That advice is out of date and harmful today.

Google draws your page the way a browser does. If it cannot load your CSS and JavaScript, it may see a broken or empty page. Google’s robots.txt guide warns not to block these files when that makes the page harder for Google to understand.

Path or URL Block it? Why
/wp-admin/ Yes Private dashboard. Nothing there should rank.
/wp-admin/admin-ajax.php No, allow it Themes and plugins use it to load public content.
/?s= internal search Yes Creates endless thin pages that waste crawling.
/wp-content/uploads/ Never Your images live here. Blocking it drops them from image search.
/wp-includes/ and theme CSS or JS Never Google needs them to draw your page correctly.
Category and tag pages Usually no If they are thin, use noindex instead, so Google can still follow links.
Pages you want gone from Google No, use noindex A blocked page cannot show Google its noindex tag.
Whole staging site Yes, plus a password Keeps test copies out of search.
Quick reference for common WordPress paths.

Your uploads folder deserves a special mention. Images bring real traffic from image search. Instead of hiding heavy images from bots, make them lighter. Our Smart Image Compressor and Bulk Image Resizer do that in the browser, and our guide on how to speed up a slow WordPress site covers the rest.

Robots.txt or noindex: which one do you need?

These two tools do different jobs, and mixing them up is one of the easiest mistakes to make on a WordPress site. Robots.txt controls crawling, whether a bot may read a page. Noindex controls indexing, whether a page may appear in results.

The trap is using both on the same page. If robots.txt blocks a page, Google never reads it, so it never sees the noindex tag. The page can stay in results for months. Use this diagram to choose:

Robots.txt or noindex?Do you want this pageOUT of Google results?YESNOUse noindexLeave it crawlable soGoogle can SEE the tag.Do not block it.Is it a junk URL thatwastes crawling, likesearch results oradmin pages?YESNODisallow it inrobots.txtSaves crawl timeLeave itopenThe common mistakeBlocking a page AND adding noindex. Google cannotcrawl it, so it never sees the noindex tag.
Robots.txt controls crawling. Noindex controls what shows in search.

Should you block AI crawlers like GPTBot and ClaudeBot?

This is the question most robots.txt guides skip, and it matters more every month. AI companies now run several different bots. Some collect pages to train models. Others fetch pages so the AI can show and cite you in its answers. Blocking the wrong one can cost you visibility.

The rules differ by company, so here is what each one says in its own documentation, in plain words:

Bot name Company What it does If you block it
GPTBot OpenAI Collects pages for model training Your content is left out of training. ChatGPT search is not affected.
OAI-SearchBot OpenAI Finds pages for ChatGPT search You stop appearing in ChatGPT search answers.
ClaudeBot Anthropic Collects pages for model training Future content is left out of training.
Claude-SearchBot Anthropic Indexes pages for Claude search Less visibility in Claude search results.
Claude-User Anthropic Fetches a page when a user asks Claude cannot read your page for that user.
Google-Extended Google Controls use in Gemini training and grounding No effect on Google Search ranking or inclusion.
Sources: OpenAI’s crawler page, Anthropic’s help article and Google’s crawler list.

For a small business that wants customers, my advice is simple. If you want to opt out of training, block the training bots: GPTBot, ClaudeBot and Google-Extended. Keep the search and user bots open. Being cited by AI assistants can send you visitors, and blocking those bots throws that chance away.

One technical detail catches people out. A bot follows only the group of rules written for its own name. If you add a group for GPTBot, it ignores everything under User-agent: *. So repeat any general rules inside a specific group if you still want them applied.

What happens if your robots.txt breaks?

A missing robots.txt is harmless. If the file returns a “not found” error, Google simply assumes there are no rules. A server error is a very different story.

If your server returns a 5xx error (a server side failure) for robots.txt, Google treats your whole site as blocked at first. Google’s robots.txt specification describes the steps:

If robots.txt returns a 5xx server errorHow Google says it behavesFirst 12 hoursGoogle stops crawling your site and retries.Up to 30 daysGoogle uses the last copy of robots.txtit saved, and keeps retrying.After 30 daysIf the rest of the site works, Google actsas if there is no robots.txt. If the siteis also down, crawling stays stopped.A 404 is different: Google treats it as “no rules”.
Source: Google Search Central robots.txt specification.

That first stage is the dangerous one. A broken security plugin, a bad server rule or a firewall that blocks Googlebot can make robots.txt fail while your pages look fine to you. Crawling quietly stops. So after any server or plugin change, open yoursite.com/robots.txt in a browser and make sure it loads.

The same specification gives a few more limits worth knowing. Google reads only the first 500 KiB of the file and ignores the rest. Paths are case sensitive, so /Shop/ and /shop/ are different. Google ignores crawl-delay completely, and it only supports four fields: user-agent, allow, disallow and sitemap. Google also keeps a saved copy for up to 24 hours, so changes are not instant.

A safe starter robots.txt for most WordPress sites

Here is a sensible file for a typical small business site or blog. Replace the domain and check your real sitemap address first. Rank Math and Yoast usually use /sitemap_index.xml, and WordPress’s own sitemap lives at /wp-sitemap.xml.

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/

# Optional: opt out of AI training only
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

Sitemap: https://yoursite.com/sitemap_index.xml

Notice what is missing. No /wp-content/ block, no /wp-includes/ block, no category block. Short is safer. Every extra line is one more chance to hide something you wanted found.

If you run a shop, add only the filter parameters you have seen causing trouble in your crawl reports. Do not copy a long list from another site. Their URLs are not your URLs.

How do you test robots.txt after changing it?

Run these checks every time you edit the file:

  • Open yoursite.com/robots.txt in a private browser window. It should load as plain text with no error.
  • In Google Search Console, open the robots.txt report. It shows when Google last fetched your file and any problems it found. You can ask Google to fetch it again after a fix.
  • Use the URL Inspection tool on your homepage and one key service page. Both should say crawling is allowed.
  • Check one image URL from /wp-content/uploads/ the same way. It should also be allowed.

If you would rather not write the file by hand, our free Robots.txt Generator lets you pick the rules and copy a clean, correctly formatted file. Then run the same checks above.

Related reading

Need help with your site or business systems?

Robots.txt is one small setting, but it sits on top of everything else your website does. I build WordPress sites, custom CRMs and business automations for small firms, and I check details like this as part of every build. If you want a site that brings in leads and a system that follows them up, get in touch here.

 

0 0 votes
Rating
0 0 votes
Rating
Subscribe
Notify of
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x