robots.txt vs noindex: which one keeps a page out of Google?
They sound similar but do very different jobs, and using the wrong one can backfire. A plain-English guide with examples of when to use each.
If you want a page to stay out of Google, it’s tempting to block it in robots.txt and move on. Unfortunately that often doesn’t work, and occasionally it makes things worse. Here’s the difference between the two tools and when to use each.
What robots.txt does
robots.txt is a small text file at the root of your domain (example.com/robots.txt) that tells crawlers which URLs they may fetch. It controls crawling, not indexing.
If a blocked URL has links pointing to it from other pages or sites, Google can still list it in search results. It just won’t know what’s on the page, so the result may appear with no description.
Use robots.txt to:
- stop crawlers wasting time on endless URL variations (filters, internal search results, calendars),
- keep crawlers out of admin or system areas,
- tell search engines where your sitemap is.
What noindex does
noindex is an instruction on the page itself, usually a meta tag in the <head>:
<meta name="robots" content="noindex">
It tells search engines: you can visit this page, but don’t show it in results. It controls indexing.
Use noindex for:
- thank-you and confirmation pages,
- thin pages like tag archives with one post,
- internal search results pages,
- staging or test pages that must stay public.
The mistake to avoid
Don’t combine them on the same page. If you block a page in robots.txt, Google can’t crawl it, which means it never sees the noindex tag. The page can stay in the index indefinitely.
To remove a page from Google: add noindex, keep the page crawlable, and wait for Google to revisit it. Only once it has dropped out should you consider blocking it in robots.txt, if you still need to.
What about private content?
Neither method protects private information. robots.txt is a public file that anyone can read, and both are polite requests that well-behaved crawlers follow. Anything confidential belongs behind a password.
Quick reference
| Goal | Use |
|---|---|
| Hide a page from search results | noindex |
| Stop crawlers visiting endless URL variations | robots.txt |
| Point search engines to your sitemap | robots.txt |
| Keep something truly private | Password protection |
Build yours
The Robots.txt Generator creates a clean file with your blocked folders and sitemap URL. For page-level tags, the Meta Tag Generator produces the head tags you need to paste into your pages.