The robots.txt file is a small text file with a big responsibility. It tells search engine crawlers which parts of your website they are allowed to visit. Used correctly, it helps crawlers focus on your important content and avoid unnecessary areas. Used incorrectly, it can accidentally block your entire website from search engines.

This guide explains how robots.txt works, how to read and write rules, which situations call for it and which common mistakes to avoid.

At a Glance: Robots.txt sits at yourdomain.com/robots.txt and controls crawling, not indexing. Use it to keep bots out of low-value areas. Never use it to hide private information or to remove pages from search results; use noindex or authentication instead.

Where Robots.txt Lives

The file must be located at the root of your domain, for example https://example.com/robots.txt. Each subdomain needs its own file. Crawlers check this file before crawling your site. WordPress generates a virtual robots.txt by default, and SEO plugins let you edit it from the dashboard.

Basic Robots.txt Syntax

— — —
User-agent Which crawler the rules apply to User-agent: * (all crawlers)
Disallow Path the crawler should not visit Disallow: /wp-admin/
Allow Exception within a disallowed path Allow: /wp-admin/admin-ajax.php
Sitemap Location of your XML sitemap Sitemap: https://example.com/sitemap_index.xml

Rules are grouped by user-agent. A blank Disallow line means nothing is blocked. Paths are case-sensitive.

A Typical WordPress Robots.txt

A safe, simple robots.txt for many WordPress sites blocks the admin area while allowing the AJAX file that some themes and plugins need, and lists the sitemap. For most small blogs, that is all you need. Additional rules should only be added when there is a clear reason.

What Robots.txt Is Good For

  • Preventing crawlers from wasting time on admin and login areas
  • Blocking internal search result pages that create endless URL variations
  • Limiting crawling of filtered or sorted URLs on large ecommerce sites
  • Keeping staging or development areas from being crawled (alongside password protection)
  • Pointing crawlers to your sitemap

What Robots.txt Cannot Do

It does not remove pages from search results

If a blocked page has links pointing to it, search engines may still index its URL without content. To keep a page out of search results, allow crawling and use a noindex meta tag, or protect the page with a password.

It does not protect private information

Robots.txt is public. Anyone can read it, and it can even reveal where sensitive areas are. Never rely on it for security.

Not every bot obeys it

Reputable search engines follow robots.txt, but malicious bots may ignore it.

Robots.txt vs Noindex

Goal Use Robots.txt? Use Noindex?
— — —
Keep a thank-you page out of results No Yes
Hide confidential content No No (use authentication)

An important point: if you block a page in robots.txt, search engines cannot see the noindex tag on it. So do not block a page you want to noindex.

Wildcards and Special Characters

Google supports two special characters:

  • Asterisk (*) matches any sequence of characters. For example, Disallow: /*?s= blocks internal search URLs.
  • Dollar sign ($) marks the end of a URL. For example, Disallow: /*.pdf$ blocks URLs ending in .pdf.

Use these carefully and test the results before relying on them.

Blocking AI Crawlers

Some website owners choose to block specific AI training crawlers by naming their user-agents in robots.txt. This is a business decision. Blocking these bots generally does not affect standard search engine crawling, but check each crawler’s documentation to understand what it is used for.

How to Test Robots.txt

  • Open yourdomain.com/robots.txt in a browser to check the live file
  • Use the robots.txt report in Google Search Console to see the version Google fetched and any errors
  • Use the URL Inspection tool to check whether a specific page is blocked
  • Test important URLs after any change

Common Robots.txt Mistakes

  • Disallow: / for all user-agents, which blocks the whole site, often left over from a staging site
  • Blocking CSS and JavaScript files that search engines need to render pages
  • Blocking pages that have noindex tags, so the tag is never seen
  • Using robots.txt to hide sensitive data
  • Typos and incorrect capitalisation in paths
  • Forgetting that each subdomain needs its own file

WordPress-Specific Tip

In Settings, then Reading, WordPress has an option called “Discourage search engines from indexing this site”. It is useful while building a site but disastrous if left enabled after launch. Always check this setting when your site goes live.

Real-World Example

After moving a website from a staging server to the live domain, a business noticed its traffic falling to almost zero within weeks. The robots.txt file still contained “Disallow: /” from the staging environment, and the WordPress discourage option was also enabled. After fixing both settings and requesting reindexing of key pages in Search Console, traffic gradually returned. A simple post-launch checklist would have prevented the problem.

Frequently Asked Questions

Do I need a robots.txt file?

It is not strictly required, but it is recommended. Without one, crawlers assume they may crawl everything.

Can robots.txt improve rankings?

Not directly. It helps crawlers spend time on important pages, which can support efficient crawling on larger sites.

How do I edit robots.txt in WordPress?

Most SEO plugins include a file editor. You can also upload a physical robots.txt file to your site’s root folder.

Should I block category or tag pages?

Usually not with robots.txt. If they are thin, use noindex instead so search engines can still crawl links on them.

How quickly do changes take effect?

Search engines usually cache robots.txt for up to about a day, so changes are typically picked up quickly.

Should I list my sitemap in robots.txt?

Yes. It is a simple way to help all crawlers find your sitemap.

Conclusion

Robots.txt is a powerful tool for guiding crawlers, but it should be used with care. Keep it simple, block only areas that truly do not need crawling, never use it for privacy or de-indexing and always test after changes. A clean robots.txt file helps search engines focus on the content that matters most.

Helpful Links

Comments are closed.