---
title: "robots.txt, meta robots and X-Robots-Tag"
description: "Right tool, right job Goal Tool Notes --- --- --- Stop crawling of a path robots.txt Disallow Does not guarantee de-indexing Keep a crawlable page out of…"
url: https://optimizeall.com/learn/technical-seo-mastery/robots-txt-meta-robots-x-robots-tag
updated: 2026-10-05
---

Technical SEO Mastery · Crawl budget, robots.txt and indexing directives · lesson 4 of 18 · 14 min

# robots.txt, meta robots and X-Robots-Tag

## Right tool, right job

| Goal | Tool | Notes |
|---|---|---|
| Stop crawling of a path | robots.txt `Disallow` | Does not guarantee de-indexing |
| Keep a crawlable page out of the index | `meta name="robots" content="noindex"` | Page must *not* be blocked in robots.txt |
| Noindex non-HTML files (PDF, images) | `X-Robots-Tag: noindex` HTTP header | Set at server/CDN level |
| Consolidate duplicates | `rel="canonical"` or redirects | Covered in Module 3 |
| Remove urgently | Search Console Removals tool | Temporary (about six months); fix the source too |
| Protect private content | Authentication | Never rely on robots.txt for secrecy — it is public |

## robots.txt fundamentals

The file lives at the root of each host and protocol: `https://www.example.com/robots.txt` does not cover `https://shop.example.com/`. It follows RFC 9309 (the Robots Exclusion Protocol standard).

```text
# robots.txt for https://www.example.com
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /search
Disallow: /*?*sessionid=
Allow: /search/help

User-agent: Googlebot-Image
Disallow: /internal-diagrams/

Sitemap: https://www.example.com/sitemap_index.xml
```

Rules to remember:

- **Grouping:** a crawler obeys only the *most specific* user-agent group that matches it. If you add a `User-agent: Googlebot` group, Googlebot ignores the `*` group entirely — repeat any shared rules.
- **Precedence:** for Google, the most specific (longest) matching rule wins; if an Allow and Disallow are equally specific, Allow wins.
- **Wildcards:** `*` matches any sequence; `$` anchors the end. `Disallow: /*.pdf$` blocks URLs ending in .pdf.
- **Case-sensitive paths:** `/Search` and `/search` are different.
- **Size:** Google processes the first 500 KiB.
- **Unsupported by Google:** `noindex`, `crawl-delay`, `nofollow` lines in robots.txt. Bing does support crawl-delay.
- **Errors:** if robots.txt returns 4xx, Google treats it as "no restrictions". If it returns 5xx for an extended period, Google may stop crawling the site. A broken robots.txt server can therefore halt crawling.

## Meta robots and X-Robots-Tag

```html
<meta name="robots" content="noindex, follow">
<meta name="googlebot" content="noindex">
<meta name="robots" content="max-snippet:160, max-image-preview:large">
```

```text
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex, nofollow
```

Useful values: `noindex`, `nofollow`, `none` (= noindex, nofollow), `nosnippet`, `max-snippet:[n]`, `max-image-preview:[none|standard|large]`, `max-video-preview:[n]`, `notranslate`, `noimageindex`, `unavailable_after:[date]`. There is also the `data-nosnippet` HTML attribute to exclude a section of a page from snippets.

Note that snippet controls such as `nosnippet` also apply to how Google may show your content in AI features in Search. Treat them as a blunt instrument — you will lose normal snippets too.

Example Apache and Nginx configuration for noindexing all PDFs:

```apache
<FilesMatch "\.pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>
```

```nginx
location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex" always;
}
```

## The classic trap: blocked + noindex

If you Disallow a path in robots.txt **and** add noindex to the pages, Google cannot crawl them, never sees the noindex, and may keep the URLs indexed from links. To de-index: allow crawling, serve noindex, wait until the URLs drop out (confirm in URL Inspection or the Page indexing report), and only then consider a Disallow if you also want to stop crawling.

## Staging and development sites

Staging environments get indexed constantly because someone links to them or a sitemap leaks. The robust approach is **HTTP authentication** (or IP allow-listing). A global `X-Robots-Tag: noindex` header is a good second layer. Do not rely on `Disallow: /` alone, and make sure deploy scripts cannot copy staging robots rules to production — a production `Disallow: /` is one of the most damaging one-line mistakes in SEO.

## Testing checklist

1. Fetch `/robots.txt` on every host (www, non-www, subdomains, each protocol) and confirm a 200 or intended 404.
2. Review the robots.txt report in Search Console for fetch errors and the last fetched version.
3. Test critical URLs (home, categories, products, articles, JS/CSS) against the rules with a crawler or parser.
4. Crawl the site and list all URLs with `noindex` in meta or header; confirm each one is intended.
5. Check that no noindexed URL is in the XML sitemap and no canonical points to a noindexed URL.
6. Add robots.txt and key directives to your release regression tests.

## AI crawlers in the same file

robots.txt now also carries decisions about AI crawlers — training bots such as `GPTBot`, `ClaudeBot`, `CCBot`, and control tokens such as `Google-Extended` and `Applebot-Extended`, alongside AI search bots such as `OAI-SearchBot`, `Claude-SearchBot` and `PerplexityBot`. The same grouping rule applies: a bot with its own group ignores the `*` group, so repeat shared disallows. Blocking `Google-Extended` does not affect Google Search or AI Overviews. The **AI Search Optimization** course covers the policy side in depth; technically, treat every AI group like any other group and include it in your regression tests.

## Hands-on: a robots.txt regression test for CI

```python
# test_robots.py — run in CI against staging and production (pytest)
import urllib.request
from urllib import robotparser
import pytest

HOST = "https://www.example.com"
MUST_ALLOW = {"Googlebot": ["/", "/products/brass-lantern/", "/static/app.js"],
              "Bingbot": ["/", "/blog/"]}
MUST_BLOCK = {"Googlebot": ["/cart/", "/search?q=test"]}

@pytest.fixture(scope="module")
def rp():
    txt = urllib.request.urlopen(HOST + "/robots.txt", timeout=10).read().decode()
    assert "Disallow: /\n" not in txt.replace("\r", "") + "\n" or "User-agent: *" not in txt, \
        "Sitewide Disallow found - did staging rules ship?"
    parser = robotparser.RobotFileParser(); parser.parse(txt.splitlines()); return parser

@pytest.mark.parametrize("agent,path", [(a, p) for a, ps in MUST_ALLOW.items() for p in ps])
def test_allowed(rp, agent, path):
    assert rp.can_fetch(agent, HOST + path), f"{agent} blocked from {path}"

@pytest.mark.parametrize("agent,path", [(a, p) for a, ps in MUST_BLOCK.items() for p in ps])
def test_blocked(rp, agent, path):
    assert not rp.can_fetch(agent, HOST + path), f"{agent} can fetch {path}"
```

`urllib.robotparser` evaluates rules in file order rather than Google's longest-match precedence, so keep test paths unambiguous and confirm edge cases in Search Console's robots.txt report. Also add a check that key templates don't output `noindex` (fetch a sample URL and assert the meta robots and `X-Robots-Tag` header).

## Worked example 2: noindex that never took effect

A Karachi e-learning site (illustrative) added `noindex` to 8,000 thin tag pages and, in the same release, disallowed `/tag/` in robots.txt "to be thorough". Three months later the pages were still indexed. Googlebot could no longer fetch them, so it never saw the noindex. The fix followed the correct sequence: remove the Disallow, keep noindex, confirm the URLs dropping out in the Page indexing report, and only then consider re-adding the Disallow to save crawling.

## Common mistakes

- Blocking CSS/JS "to save budget", breaking rendering.
- Assuming robots.txt hides confidential URLs — it advertises them.
- Leaving `noindex` on templates after launch.
- Forgetting that a specific Googlebot group overrides the `*` group.

## Video lecture: robots.txt, meta robots and X-Robots-Tag

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. robots.txt, meta robots, X-Robots-Tag
2. Right tool, right job
3. robots.txt fundamentals
4. Facts people forget
5. Index directives
6. The blocked + noindex trap
7. Staging sites
8. AI groups + hands-on tests
9. Example 1: blocked CSS and JS
10. Three layers of robots testing
11. Watch me do it: robots regression tests
12. Example 2: noindex that never worked (illustrative)
13. Mistakes, recap, try this now

## Lecture transcript

### robots.txt, meta robots, X-Robots-Tag

robots.txt, meta robots and the X-Robots-Tag header look simple. They're also behind some of the most expensive mistakes in SEO, including the one-line production Disallow that removes a whole site from crawling. In this lecture you'll learn which tool does which job, the rules that trip up experts, the classic blocked-plus-noindex trap, and how to put robots.txt into automated tests so a bad deploy can't slip through.

### Right tool, right job

Start with the job, not the tool. To stop crawling of a path, use robots.txt Disallow, but it doesn't guarantee de-indexing. To keep a crawlable page out of the index, use a meta robots noindex, and don't block it in robots.txt. For non-HTML files like PDFs, use the X-Robots-Tag HTTP header. To consolidate duplicates, use canonicals or redirects. For an urgent temporary removal, use Search Console's Removals tool, which lasts about six months. And to protect private content, use authentication, because robots.txt is public and advertises your paths.

### robots.txt fundamentals

Now robots.txt fundamentals. It lives at the root of each host and protocol, so www and a shop subdomain each need their own. It follows RFC ninety-three-oh-nine. Four rules matter most. Grouping: a crawler obeys only the most specific user-agent group that matches it, so a Googlebot group means Googlebot ignores the star group entirely. Precedence: for Google, the longest matching rule wins, and if Allow and Disallow tie, Allow wins. Wildcards: star matches any sequence and dollar anchors the end. And paths are case-sensitive.

### Facts people forget

A few more facts people forget. Google processes only the first five hundred kibibytes of robots.txt. Google doesn't support noindex, crawl-delay or nofollow lines in robots.txt, although Bing does support crawl-delay. And error handling matters. If robots.txt returns a four-hundred-level error, Google treats it as no restrictions. If it returns five-hundred-level errors for an extended period, Google may stop crawling the site. So a broken robots.txt server can quietly halt crawling.

### Index directives

Meta robots and X-Robots-Tag carry the index directives. Common values: noindex, nofollow, none, nosnippet, max-snippet, max-image-preview, max-video-preview, noimageindex and unavailable after. There's also the data-nosnippet attribute to exclude a section from snippets. Snippet controls also apply to how Google may show your content in AI features, but they're blunt, because you lose normal snippets too. Since twenty twenty-six, Google also has a Search Console setting that excludes a site from AI Overviews and AI Mode without touching snippets. For PDFs, the lesson text has Apache and Nginx examples for an X-Robots-Tag noindex header.

### The blocked + noindex trap

Now the classic trap. You Disallow a path in robots.txt and add noindex to the pages. Google can't crawl them, so it never sees the noindex, and it may keep the URLs indexed from links. The correct sequence is: allow crawling, serve noindex, wait until the URLs drop out, confirmed in URL Inspection or the Page indexing report, and only then consider a Disallow if you also want to stop crawling. Order matters.

### Staging sites

Staging sites deserve their own warning. They get indexed constantly because someone links to them or a sitemap leaks. The robust approach is HTTP authentication or IP allow-listing, with a global X-Robots-Tag noindex header as a second layer. Don't rely on Disallow slash alone. And make sure deploy scripts can't copy staging robots rules to production, because a production Disallow slash is one of the most damaging one-line mistakes in SEO.

### AI groups + hands-on tests

Robots.txt now carries AI crawler decisions too. Training bots like GPTBot and ClaudeBot, control tokens like Google-Extended and Applebot-Extended, and AI search bots like OAI-SearchBot and PerplexityBot. Technically, treat each AI group like any other: it ignores the star group, so repeat shared disallows, and include it in your tests. The policy side is covered in the AI Search Optimization course. Now the hands-on. The lesson text includes a pytest file that fetches robots.txt, fails the build if a sitewide Disallow appears, and asserts which agents can and can't fetch key paths.

### Example 1: blocked CSS and JS

Worked example one, simple. A Birmingham law firm's developer blocks slash wp-content to hide uploads. That also blocks the theme's CSS and JavaScript, and URL Inspection shows an unstyled, broken render. The fix is removing that rule and, if some files must stay private, putting them behind authentication instead. Five minutes, and the page renders properly again.

### Three layers of robots testing

Let's make testing concrete with Search Console. The robots.txt report shows which robots.txt files Google found for the top hosts in your property, when each was last fetched, and any fetch errors or parsing warnings. After any robots.txt change, check that Google has fetched the new version before you expect crawl behaviour to change. For individual URLs, URL Inspection tells you whether crawling is allowed. Combine those with your own automated tests and you have three layers: your CI test before deploy, Google's fetched version after deploy, and log evidence of how Googlebot actually behaves.

### Watch me do it: robots regression tests

Watch me do it. I'll add robots regression tests to a real project. Step one: I list what must always be allowed for Googlebot: the home page, one product, one category, and the main JavaScript bundle. And what must be blocked: the cart and internal search. Step two: I copy the pytest file from the lesson into the tests folder and fill in those paths. Step three: I run pytest against production. Seven pass. One fails: Googlebot blocked from slash static slash app dot js. I open robots.txt and find a Disallow for slash static, added last year to hide design files. That rule is also blocking the JavaScript needed for rendering. Step four: I check URL Inspection for a product page, test live, and open the page resources. The app bundle shows as blocked, and the screenshot is half-rendered. Confirmed. Step five: I change the rule to disallow only the design subfolder, run the tests again, and everything passes. Step six: after deploy, I open Search Console's robots.txt report and confirm Google fetched the new version. And now the test runs in CI on every deploy, so the next accidental block fails the build instead of reaching production.

### Example 2: noindex that never worked (illustrative)

Worked example two, with illustrative details. A Karachi e-learning site added noindex to eight thousand thin tag pages and, in the same release, disallowed the tag folder in robots.txt to be thorough. Three months later, the pages were still indexed. Googlebot couldn't fetch them, so it never saw the noindex. The fix followed the correct sequence: remove the Disallow, keep noindex, confirm the URLs dropping out in Page indexing, and only then consider re-adding the Disallow.

### Mistakes, recap, try this now

Common mistakes: blocking CSS and JavaScript to save budget, which breaks rendering. Assuming robots.txt hides confidential URLs, when it advertises them. Leaving noindex on templates after launch. And forgetting that a specific Googlebot group overrides the star group. Recap: use each tool for its job, respect grouping and precedence, never combine a block with a noindex you need Google to see, and protect staging properly. Try this now. Add the pytest file to your project, fill in five must-allow and three must-block paths, and run it against production today.

## Key takeaways

- robots.txt manages crawling; noindex (meta or X-Robots-Tag) manages indexing — never combine Disallow with noindex when de-indexing.
- A crawler follows only the most specific matching user-agent group; Google applies the longest matching rule.
- Use X-Robots-Tag for PDFs and other non-HTML files.
- Protect staging with authentication, not robots.txt, and regression-test robots rules on every release.

## Try it

Audit a live robots.txt against the testing checklist and write a corrected version with comments explaining each rule.

- [Previous: Crawl budget and log-file analysis](https://optimizeall.com/learn/technical-seo-mastery/crawl-budget-and-log-file-analysis)
- [Next: Beyond Google: Bing Webmaster Tools, IndexNow and AI-era crawling](https://optimizeall.com/learn/technical-seo-mastery/bing-webmaster-tools-and-indexnow)
- [All lessons of Technical SEO Mastery](https://optimizeall.com/learn/technical-seo-mastery)
