Technical SEO MasteryCrawl budget, robots.txt and indexing directives · Lesson 4 of 18

robots.txt, meta robots and X-Robots-Tag

Article · 14 min · 8 min lecture

Video lecture

robots.txt, meta robots and X-Robots-Tag

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

robots.txt, meta robots, X-Robots-Tag

  • The right tool for each job
  • Rules that trip up experts
  • Automated regression tests

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Right tool, right job

GoalToolNotes
Stop crawling of a pathrobots.txt DisallowDoes not guarantee de-indexing
Keep a crawlable page out of the indexmeta name="robots" content="noindex"Page must not be blocked in robots.txt
Noindex non-HTML files (PDF, images)X-Robots-Tag: noindex HTTP headerSet at server/CDN level
Consolidate duplicatesrel="canonical" or redirectsCovered in Module 3
Remove urgentlySearch Console Removals toolTemporary (about six months); fix the source too
Protect private contentAuthenticationNever rely on robots.txt for secrecy — it is public

robots.txt fundamentals

The file lives at the root of each host and protocol: https://www.example.com/robots.txt does not cover https://shop.example.com/. It follows RFC 9309 (the Robots Exclusion Protocol standard).

# robots.txt for https://www.example.com
User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /search
Disallow: /*?*sessionid=
Allow: /search/help

User-agent: Googlebot-Image
Disallow: /internal-diagrams/

Sitemap: https://www.example.com/sitemap_index.xml

Rules to remember:

  • Grouping: a crawler obeys only the most specific user-agent group that matches it. If you add a User-agent: Googlebot group, Googlebot ignores the * group entirely — repeat any shared rules.
  • Precedence: for Google, the most specific (longest) matching rule wins; if an Allow and Disallow are equally specific, Allow wins.
  • Wildcards: * matches any sequence; $ anchors the end. Disallow: /*.pdf$ blocks URLs ending in .pdf.
  • Case-sensitive paths: /Search and /search are different.
  • Size: Google processes the first 500 KiB.
  • Unsupported by Google: noindex, crawl-delay, nofollow lines in robots.txt. Bing does support crawl-delay.
  • Errors: if robots.txt returns 4xx, Google treats it as "no restrictions". If it returns 5xx for an extended period, Google may stop crawling the site. A broken robots.txt server can therefore halt crawling.

Meta robots and X-Robots-Tag

<meta name="robots" content="noindex, follow">
<meta name="googlebot" content="noindex">
<meta name="robots" content="max-snippet:160, max-image-preview:large">
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex, nofollow

Useful values: noindex, nofollow, none (= noindex, nofollow), nosnippet, max-snippet:[n], max-image-preview:[none|standard|large], max-video-preview:[n], notranslate, noimageindex, unavailable_after:[date]. There is also the data-nosnippet HTML attribute to exclude a section of a page from snippets.

Note that snippet controls such as nosnippet also apply to how Google may show your content in AI features in Search. Treat them as a blunt instrument — you will lose normal snippets too.

Example Apache and Nginx configuration for noindexing all PDFs:

<FilesMatch "\.pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>
location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex" always;
}

The classic trap: blocked + noindex

If you Disallow a path in robots.txt and add noindex to the pages, Google cannot crawl them, never sees the noindex, and may keep the URLs indexed from links. To de-index: allow crawling, serve noindex, wait until the URLs drop out (confirm in URL Inspection or the Page indexing report), and only then consider a Disallow if you also want to stop crawling.

Staging and development sites

Staging environments get indexed constantly because someone links to them or a sitemap leaks. The robust approach is HTTP authentication (or IP allow-listing). A global X-Robots-Tag: noindex header is a good second layer. Do not rely on Disallow: / alone, and make sure deploy scripts cannot copy staging robots rules to production — a production Disallow: / is one of the most damaging one-line mistakes in SEO.

Testing checklist

  1. Fetch /robots.txt on every host (www, non-www, subdomains, each protocol) and confirm a 200 or intended 404.
  2. Review the robots.txt report in Search Console for fetch errors and the last fetched version.
  3. Test critical URLs (home, categories, products, articles, JS/CSS) against the rules with a crawler or parser.
  4. Crawl the site and list all URLs with noindex in meta or header; confirm each one is intended.
  5. Check that no noindexed URL is in the XML sitemap and no canonical points to a noindexed URL.
  6. Add robots.txt and key directives to your release regression tests.

AI crawlers in the same file

robots.txt now also carries decisions about AI crawlers — training bots such as GPTBot, ClaudeBot, CCBot, and control tokens such as Google-Extended and Applebot-Extended, alongside AI search bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. The same grouping rule applies: a bot with its own group ignores the * group, so repeat shared disallows. Blocking Google-Extended does not affect Google Search or AI Overviews. The AI Search Optimization course covers the policy side in depth; technically, treat every AI group like any other group and include it in your regression tests.

Hands-on: a robots.txt regression test for CI

# test_robots.py — run in CI against staging and production (pytest)
import urllib.request
from urllib import robotparser
import pytest

HOST = "https://www.example.com"
MUST_ALLOW = {"Googlebot": ["/", "/products/brass-lantern/", "/static/app.js"],
              "Bingbot": ["/", "/blog/"]}
MUST_BLOCK = {"Googlebot": ["/cart/", "/search?q=test"]}

@pytest.fixture(scope="module")
def rp():
    txt = urllib.request.urlopen(HOST + "/robots.txt", timeout=10).read().decode()
    assert "Disallow: /\n" not in txt.replace("\r", "") + "\n" or "User-agent: *" not in txt, \
        "Sitewide Disallow found - did staging rules ship?"
    parser = robotparser.RobotFileParser(); parser.parse(txt.splitlines()); return parser

@pytest.mark.parametrize("agent,path", [(a, p) for a, ps in MUST_ALLOW.items() for p in ps])
def test_allowed(rp, agent, path):
    assert rp.can_fetch(agent, HOST + path), f"{agent} blocked from {path}"

@pytest.mark.parametrize("agent,path", [(a, p) for a, ps in MUST_BLOCK.items() for p in ps])
def test_blocked(rp, agent, path):
    assert not rp.can_fetch(agent, HOST + path), f"{agent} can fetch {path}"

urllib.robotparser evaluates rules in file order rather than Google's longest-match precedence, so keep test paths unambiguous and confirm edge cases in Search Console's robots.txt report. Also add a check that key templates don't output noindex (fetch a sample URL and assert the meta robots and X-Robots-Tag header).

Worked example 2: noindex that never took effect

A Karachi e-learning site (illustrative) added noindex to 8,000 thin tag pages and, in the same release, disallowed /tag/ in robots.txt "to be thorough". Three months later the pages were still indexed. Googlebot could no longer fetch them, so it never saw the noindex. The fix followed the correct sequence: remove the Disallow, keep noindex, confirm the URLs dropping out in the Page indexing report, and only then consider re-adding the Disallow to save crawling.

Common mistakes

  • Blocking CSS/JS "to save budget", breaking rendering.
  • Assuming robots.txt hides confidential URLs — it advertises them.
  • Leaving noindex on templates after launch.
  • Forgetting that a specific Googlebot group overrides the * group.

Key takeaways

  • robots.txt manages crawling; noindex (meta or X-Robots-Tag) manages indexing — never combine Disallow with noindex when de-indexing.
  • A crawler follows only the most specific matching user-agent group; Google applies the longest matching rule.
  • Use X-Robots-Tag for PDFs and other non-HTML files.
  • Protect staging with authentication, not robots.txt, and regression-test robots rules on every release.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A robots.txt has a `User-agent: *` group with Disallow: /cart/ and a separate `User-agent: Googlebot` group with only Disallow: /tmp/. Can Googlebot crawl /cart/?
  2. How should you keep 5,000 PDF files out of Google's index?
  3. What happens if robots.txt consistently returns a 5xx error?

Put it into practice

Audit a live robots.txt against the testing checklist and write a corrected version with comments explaining each rule.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.