The robots.txt file is a small text file with a big responsibility. It tells search engine crawlers which parts of your website they are allowed to visit. Used correctly, it helps crawlers focus on your important content and avoid unnecessary areas. Used incorrectly, it can accidentally block your entire website from search engines.
This guide explains how robots.txt works, how to read and write rules, which situations call for it and which common mistakes to avoid.
Where Robots.txt Lives
The file must be located at the root of your domain, for example https://example.com/robots.txt. Each subdomain needs its own file. Crawlers check this file before crawling your site. WordPress generates a virtual robots.txt by default, and SEO plugins let you edit it from the dashboard.
Basic Robots.txt Syntax
| — | — | — |
|---|---|---|
| User-agent | Which crawler the rules apply to | User-agent: * (all crawlers) |
| Disallow | Path the crawler should not visit | Disallow: /wp-admin/ |
| Allow | Exception within a disallowed path | Allow: /wp-admin/admin-ajax.php |
| Sitemap | Location of your XML sitemap | Sitemap: https://example.com/sitemap_index.xml |
Rules are grouped by user-agent. A blank Disallow line means nothing is blocked. Paths are case-sensitive.
A Typical WordPress Robots.txt
A safe, simple robots.txt for many WordPress sites blocks the admin area while allowing the AJAX file that some themes and plugins need, and lists the sitemap. For most small blogs, that is all you need. Additional rules should only be added when there is a clear reason.
What Robots.txt Is Good For
- Preventing crawlers from wasting time on admin and login areas
- Blocking internal search result pages that create endless URL variations
- Limiting crawling of filtered or sorted URLs on large ecommerce sites
- Keeping staging or development areas from being crawled (alongside password protection)
- Pointing crawlers to your sitemap
What Robots.txt Cannot Do
It does not remove pages from search results
If a blocked page has links pointing to it, search engines may still index its URL without content. To keep a page out of search results, allow crawling and use a noindex meta tag, or protect the page with a password.
It does not protect private information
Robots.txt is public. Anyone can read it, and it can even reveal where sensitive areas are. Never rely on it for security.
Not every bot obeys it
Reputable search engines follow robots.txt, but malicious bots may ignore it.
Robots.txt vs Noindex
| Goal | Use Robots.txt? | Use Noindex? |
|---|---|---|
| — | — | — |
| Keep a thank-you page out of results | No | Yes |
| Hide confidential content | No | No (use authentication) |
An important point: if you block a page in robots.txt, search engines cannot see the noindex tag on it. So do not block a page you want to noindex.
Wildcards and Special Characters
Google supports two special characters:
- Asterisk (*) matches any sequence of characters. For example, Disallow: /*?s= blocks internal search URLs.
- Dollar sign ($) marks the end of a URL. For example, Disallow: /*.pdf$ blocks URLs ending in .pdf.
Use these carefully and test the results before relying on them.
Blocking AI Crawlers
Some website owners choose to block specific AI training crawlers by naming their user-agents in robots.txt. This is a business decision. Blocking these bots generally does not affect standard search engine crawling, but check each crawler’s documentation to understand what it is used for.
How to Test Robots.txt
- Open yourdomain.com/robots.txt in a browser to check the live file
- Use the robots.txt report in Google Search Console to see the version Google fetched and any errors
- Use the URL Inspection tool to check whether a specific page is blocked
- Test important URLs after any change
Common Robots.txt Mistakes
- Disallow: / for all user-agents, which blocks the whole site, often left over from a staging site
- Blocking CSS and JavaScript files that search engines need to render pages
- Blocking pages that have noindex tags, so the tag is never seen
- Using robots.txt to hide sensitive data
- Typos and incorrect capitalisation in paths
- Forgetting that each subdomain needs its own file
WordPress-Specific Tip
In Settings, then Reading, WordPress has an option called “Discourage search engines from indexing this site”. It is useful while building a site but disastrous if left enabled after launch. Always check this setting when your site goes live.
Real-World Example
After moving a website from a staging server to the live domain, a business noticed its traffic falling to almost zero within weeks. The robots.txt file still contained “Disallow: /” from the staging environment, and the WordPress discourage option was also enabled. After fixing both settings and requesting reindexing of key pages in Search Console, traffic gradually returned. A simple post-launch checklist would have prevented the problem.
Frequently Asked Questions
Do I need a robots.txt file?
It is not strictly required, but it is recommended. Without one, crawlers assume they may crawl everything.
Can robots.txt improve rankings?
Not directly. It helps crawlers spend time on important pages, which can support efficient crawling on larger sites.
How do I edit robots.txt in WordPress?
Most SEO plugins include a file editor. You can also upload a physical robots.txt file to your site’s root folder.
Should I block category or tag pages?
Usually not with robots.txt. If they are thin, use noindex instead so search engines can still crawl links on them.
How quickly do changes take effect?
Search engines usually cache robots.txt for up to about a day, so changes are typically picked up quickly.
Should I list my sitemap in robots.txt?
Yes. It is a simple way to help all crawlers find your sitemap.
Conclusion
Robots.txt is a powerful tool for guiding crawlers, but it should be used with care. Keep it simple, block only areas that truly do not need crawling, never use it for privacy or de-indexing and always test after changes. A clean robots.txt file helps search engines focus on the content that matters most.