Table of Contents
Here’s the short version: robots.txt is a plain text file at the root of your site that politely asks search engine crawlers which parts of your site to skip. Learning how to use robots.txt well means blocking the genuinely useless stuff (cart pages, internal search results, parameter chaos), never blocking the stuff crawlers need (your CSS, your JavaScript, your important pages), and understanding two truths that trip up almost everyone: robots.txt is not a security tool, and it is not a reliable way to remove pages from Google. For deindexing, you need a noindex tag — and the page has to stay crawlable for Google to see it.
Okay, deep breath. I know “technical SEO” can feel like the part of the house where the wiring lives and you’d rather not open the panel. But I promise you this: robots.txt is one of the most learnable files in all of SEO. It’s short. It’s human-readable. And once you understand the handful of rules — and the two or three famous ways people hurt themselves with it — you’ll be more careful with this file than plenty of folks who do this for a living.
Quick answer: how to use robots txt
- Robots.txt lives at yoursite.com/robots.txt and gives crawl instructions: User-agent (who), Disallow/Allow (what), and a Sitemap line (where your map lives).
- It’s a politeness request, not a lock. Good bots honor it; bad bots ignore it. Never use it to “hide” sensitive URLs — the file is public, so listing them is a signpost.
- It does not deindex pages. Blocked URLs can still appear in search if other sites link to them. To remove a page from Google, allow crawling and add a noindex tag.
- Block low-value crawl paths (cart, checkout, internal search results, messy parameters). Never block CSS or JavaScript — Google needs them to render your pages.
- Test every change before it ships. One stray Disallow: / can block your entire site.
What is robots.txt, really?
Robots.txt is part of something called the Robots Exclusion Protocol — a decades-old convention where well-behaved crawlers check a text file at the root of your domain before they start fetching pages. Googlebot, Bingbot, and the other major search crawlers all request yoursite.com/robots.txt first, read the rules, and crawl accordingly.
Notice the word I keep leaning on: crawl. Robots.txt governs crawling — the act of a bot fetching your pages. It does not govern indexing (whether a URL appears in search results), it does not govern ranking, and it absolutely does not govern access. Keep that distinction pinned to the front of your brain, because nearly every robots.txt disaster I’ve seen comes from blurring it.
And here’s the part nobody tells you early enough: robots.txt is a request, not a rule with teeth. Reputable crawlers honor it because it’s in everyone’s interest. Scrapers, spam bots, and anything built to misbehave will cheerfully ignore it. That single fact shapes everything about how you should — and shouldn’t — use this file.
Is robots.txt a security tool? (No. Please, no.)
This deserves its own section because the mistake is so common and so costly. Robots.txt is not security. It can never be security. Two reasons:
- The file is completely public. Anyone — your competitor, a curious reader, an attacker — can type yoursite.com/robots.txt into a browser and read every line. Right now. No login, no tools.
- Bad actors don’t obey it. The only bots that respect robots.txt are the polite ones. The bots you’d actually want to keep away from sensitive areas treat your disallow lines as suggestions at best.
Put those together and you get the painful irony: listing a sensitive path in robots.txt doesn’t hide it — it advertises it. If you write Disallow: /secret-admin-panel/, you’ve just published a signpost that says “interesting things live here” to exactly the audience you were trying to avoid. Security researchers literally check robots.txt files as a first step when probing a site, because so many people make this mistake.
So what do you do with genuinely sensitive areas? Protect them properly: authentication, password protection, IP restrictions, or simply keeping them off the public web entirely. A login wall stops crawlers and humans alike. Robots.txt stops neither — it only asks nicely. Use it to manage crawling of pages that are harmless but useless in search, never to guard anything that would hurt you if someone found it.
Can robots.txt remove a page from Google?
Here it is — the single most important lesson in this whole guide, the one I’d tattoo on the inside of every site owner’s eyelids if I could. Robots.txt does not reliably remove pages from search results.
Here’s why. When you block a URL in robots.txt, Google stops crawling it — but Google can still index it if other pages link to that URL. The search engine knows the page exists (links told it so), it just can’t fetch the content. So the URL can appear in results as a bare link, sometimes with a note like “No information is available for this page.” Blocked, yet indexed. The worst of both worlds.
The fix lives in a different tool entirely: the noindex directive, delivered through a meta robots tag in the page’s HTML or an X-Robots-Tag HTTP header. Noindex says “you may crawl me, but don’t show me in results” — and Google honors it reliably. But here’s the trap, and I need you to really hear this:
For noindex to work, the page must stay crawlable. Google has to fetch the page to see the noindex tag. If you block the page in robots.txt and add noindex, the crawler never reads your noindex — the robots.txt block stops it at the door — and the URL can linger in the index indefinitely.
So the correct deindexing recipe is the opposite of what intuition suggests:
- To remove a page from search: leave it crawlable in robots.txt, add a noindex tag, and wait for Google to recrawl it and process the removal. (For urgent cases, Search Console’s removal tool can temporarily hide it faster.)
- To save crawl attention on pages that don’t matter: block them in robots.txt — but accept that blocking is about crawling, not a guarantee they’ll never appear in results.
- Never combine robots.txt blocking with noindex on the same URL and expect deindexing. The block cancels the noindex.
If you only remember one thing from this article, make it that. Robots.txt controls crawling. Noindex controls indexing. They are different levers, and crossing them is the most common self-inflicted wound in technical SEO.
How do you write robots.txt syntax in plain English?
Good news: the entire syntax fits on a sticky note. A robots.txt file is a series of groups, and each group has two parts — who the rules apply to, and what the rules are.
User-agent: who you’re talking to
Every group starts with a User-agent line naming a crawler. User-agent: * means “everyone.” User-agent: Googlebot means “just Google’s main crawler.” Most sites only ever need the asterisk, but you can write separate groups for specific bots when you want different rules for different visitors. One wrinkle worth knowing: a specific crawler follows the most specific group that matches it and ignores the generic one — so if you create a Googlebot group, Googlebot reads only that group, not your * rules.
Disallow: what to skip
Disallow: /cart/ asks matching crawlers not to fetch any URL whose path starts with /cart/. The matching is “starts with,” so /cart/item-123 and /cart/checkout are both covered. An empty Disallow line (Disallow: with nothing after it) means “nothing is blocked” — which is a perfectly valid way to say “crawl everything.”
Allow: the exception maker
Allow carves exceptions out of a Disallow. Say you block /resources/ but want one PDF crawled: Disallow: /resources/ followed by Allow: /resources/annual-guide.pdf does exactly that. When Allow and Disallow rules conflict, the more specific (longer) rule generally wins.
Sitemap: the helpful signpost
The Sitemap line — Sitemap: https://yoursite.com/sitemap.xml — tells crawlers where your XML sitemap lives, using the full URL. It’s the one line in the file that’s an invitation rather than a restriction, and every robots.txt should have it. If you haven’t built that map yet, my guide on how to create an XML sitemap walks through it step by step — the sitemap and robots.txt are partners, and they work best configured together.
Wildcards and the dollar sign — with an honesty caveat
Two pattern characters give you finer control. The asterisk * matches any sequence of characters: Disallow: /*?sort= blocks any URL containing ?sort= anywhere in its path. The dollar sign $ anchors a match to the end of the URL: Disallow: /*.pdf$ blocks URLs that end in .pdf and nothing else.
Now the honest caveat: wildcard and $ support is an extension that major search engines like Google and Bing honor, but it isn’t guaranteed across every crawler, and the finer parsing details have shifted over the years. Before you build anything clever with patterns, verify current behavior in Google’s own robots.txt documentation and test your rules against real URLs. Crawler behavior evolves; a guide you read two years ago (including this one!) is a starting point, not gospel.
What should you actually block in robots.txt?
Knowing how to use robots.txt is mostly knowing what deserves blocking — and the honest answer is: less than you’d think. The good candidates share a profile: pages that are legitimate for users but worthless (or actively confusing) for search engines to crawl in bulk.
- Cart and checkout paths. /cart/, /checkout/, /order-confirmation/ — these are personal, session-driven, and have zero search value. Classic, safe blocks.
- Internal search results. Your site’s own search pages (/search?q=… or /?s=…) can generate effectively infinite thin URLs. Blocking them keeps crawlers focused on real content.
- Parameter chaos. Faceted navigation and tracking parameters can multiply one product page into hundreds of URL variants (?color=blue&sort=price&view=grid…). Pattern-blocking the worst offenders tames the sprawl — though pair this with canonical tags, which address the duplication at the indexing level.
- Staging and development areas — with a big asterisk. If a staging copy of your site is publicly reachable, blocking it in robots.txt is better than nothing. But the genuinely right answer is authentication: put staging behind a login. That actually prevents access and crawling, instead of politely requesting one of the two. (And remember: a disallowed staging URL can still get indexed from links.)
- Admin paths — with the security caveat ringing in your ears. Blocking /wp-admin/ or similar is common and reasonable for crawl tidiness, since those URLs should be behind a login anyway. Just be clear with yourself: the login is the security; the robots.txt line is housekeeping. Never list a “secret” admin path that isn’t properly protected, because now it isn’t secret.
Before you block anything, it’s worth knowing how crawlers currently move through your site — blocking is a scalpel, not a mood. A proper technical SEO audit will show you where crawl activity is actually being wasted, so your disallow rules solve real problems instead of imagined ones.
What should you never block in robots.txt?
This list is shorter but the stakes are higher. Here’s where the horror stories live.
Never block CSS and JavaScript
Google doesn’t just read your HTML — it renders your pages, like a browser does, to understand layout, content, and mobile-friendliness. To render, it needs your stylesheets and scripts. Block /css/ or /js/ or your theme’s asset folders and Google sees a broken, unstyled, possibly empty page — and judges your content accordingly. Years ago, blocking asset folders was common advice; today it’s actively harmful. If an old robots.txt on your site still disallows asset paths, freeing them is one of the easiest technical wins available.
Never block pages you want deindexed
You know this one now — it’s the trap from earlier, and I’m repeating it on purpose because it’s that common. Blocking a page you want removed from search prevents Google from seeing the noindex tag that would actually remove it. Deindex first (noindex, stay crawlable), and only consider blocking after the page is long gone from results, if ever.
Never block everything by accident
Let me tell you about the two most expensive characters in SEO: Disallow: /. A single slash after Disallow under User-agent: * asks every crawler to skip your entire site. It’s a sensible line on a staging server — and a catastrophe when that staging robots.txt ships to production, which is exactly how this disaster usually happens. A site migration goes live on launch day, everyone’s celebrating, and three days later someone notices organic traffic falling off a cliff. The culprit: the staging crawl-block came along for the ride.
The fix is boring and beautiful: make “check robots.txt in production” a mandatory line on your deployment checklist. Thirty seconds of reading one small file, every launch, forever. The teams that do this never star in the horror story.
Should you block AI crawlers in robots.txt?
Newer question, worth addressing honestly. A growing roster of AI-related crawlers now check robots.txt — user-agents associated with training data collection and with AI search and assistant products. Because robots.txt works on politeness, the reputable AI companies’ crawlers generally respect disallow rules aimed at them.
Two honest framings before you add rules:
- This is a business decision, not an SEO trick. Blocking AI crawlers won’t improve your Google rankings, and allowing them won’t hurt. The real question is about visibility tradeoffs: blocking training crawlers limits how your content feeds AI models, while blocking AI search crawlers may reduce your presence in AI-generated answers — which are becoming a real discovery channel. There’s no universally right answer; it depends on how you feel about each tradeoff.
- Verify current user-agent names before you write rules. The roster of AI crawlers — their names, their purposes, which company operates which — has changed repeatedly and will keep changing. Any specific list I gave you here would age badly. Check each AI company’s current crawler documentation for the exact user-agent strings, and review your rules every so often.
And the usual caveat applies double here: robots.txt only governs the crawlers that choose to honor it. It’s a published preference, not an enforcement mechanism.
How do you test robots.txt changes before they go live?
Robots.txt deserves the same respect as production code, because functionally that’s what it is — a tiny config file with site-wide blast radius. Here’s a sane testing routine:
- Use a robots.txt tester before deploying. Google Search Console provides robots.txt reporting and testing capability, and several reputable standalone testers let you paste rules and check specific URLs against them. Feed in your most important URLs and confirm they’re allowed; feed in the URLs you meant to block and confirm they’re blocked.
- Test the weird ones. Don’t just test your homepage. Test a CSS file, a JavaScript file, an image, a paginated URL, a parameter URL. Pattern rules love to match more than you intended.
- Stage the rollout when stakes are high. Changing rules on a large site? Ship the narrow version first, watch crawl behavior in Search Console’s crawl stats for a week or two, then widen if all’s well. Rules are easy to loosen and slow to recover from.
- Watch Search Console after every change. The Page indexing report will surface “Blocked by robots.txt” and the telltale “Indexed, though blocked by robots.txt” (that’s the blocked-yet-indexed state you now know to fix with noindex, not more blocking). If new crawl problems appear after a robots.txt change, my walkthrough on how to fix crawl errors covers diagnosing exactly what went sideways.
Your pre-deploy robots.txt checklist
Pin this somewhere your deployment process can’t miss it:
- ☐ No stray Disallow: / under User-agent: * (unless you truly intend to block the whole site)
- ☐ CSS, JavaScript, and image paths are all crawlable
- ☐ Nothing sensitive is listed (the file is public — blocking ≠ hiding)
- ☐ No page you’re trying to deindex is blocked (noindex needs crawl access)
- ☐ Sitemap line present, with the full absolute URL
- ☐ Key URLs tested against the rules in a robots.txt tester
- ☐ The production file — not the staging file — is what actually shipped
- ☐ A dated copy of the previous version is saved, so you can roll back in seconds
Does robots.txt help with crawl budget?
Time for some honesty that saves you from over-engineering. “Crawl budget” — the amount of crawling attention a search engine allocates to your site — is a real concept, but it’s mostly a big-site concern. If your site has hundreds of pages, or even a few thousand, Google can comfortably crawl all of it, and fiddling with robots.txt to “optimize crawl budget” will not move your rankings. I’ve watched small-site owners spend anxious weekends on this. Please don’t be them.
Where crawl management genuinely matters: very large sites — think huge e-commerce catalogs with faceted navigation, massive publishers, sites with URL parameters breeding millions of near-duplicate pages. There, blocking infinite low-value URL spaces helps crawlers spend their visits on pages that matter. If that’s you, robots.txt is one lever among several (canonicals, internal linking, and site architecture being the others). If that’s not you, write a clean simple file and spend your energy on content. That’s the honest math.
What are the limits of robots.txt?
A few boundary rules that surprise people:
- It’s per-subdomain and per-protocol. The file at yoursite.com/robots.txt governs only yoursite.com. Your blog at blog.yoursite.com needs its own file; so does shop.yoursite.com. Each host, its own robots.txt.
- It must live at the root. Crawlers look at /robots.txt exactly. A file at /pages/robots.txt is decoration.
- Paths are case-sensitive. Disallow: /Private/ does not block /private/. If your URLs vary in casing (please fix that separately), your rules need to match the real paths exactly.
- It can’t force crawling. Robots.txt can ask crawlers to skip things, but an Allow line can’t compel anyone to crawl more. Getting crawled more comes from site quality, fresh content, and good internal links.
- Not every directive you’ve seen is real. You’ll find old files containing crawl-delay and other directives with inconsistent support — Google, notably, ignores crawl-delay. When in doubt, check current documentation rather than copying a file you found somewhere.
A sensible default robots.txt template
For a typical small-to-medium site, this is a clean, safe starting point:
| Line | What it does |
|---|---|
| User-agent: * | These rules apply to all crawlers |
| Disallow: /cart/ | Skip cart pages (no search value) |
| Disallow: /checkout/ | Skip checkout flow |
| Disallow: /search/ | Skip internal search results (adjust to your site’s search path) |
| Disallow: /wp-admin/ | Skip admin paths (housekeeping — the login is the actual security) |
| Allow: /wp-admin/admin-ajax.php | WordPress exception: front-end features use this file |
| Sitemap: https://yoursite.com/sitemap.xml | Point crawlers to your XML sitemap (full URL) |
Adjust paths to match your actual site — that /search/ line, for instance, depends on how your platform structures search URLs — and delete anything that doesn’t apply. A shorter file you understand beats a longer file you copied. And notice everything this template doesn’t do: no asset blocking, no “secret” paths, no attempt to deindex anything. That restraint is the point.
Common robots.txt mistakes (and how to recover)
| Mistake | What happens | The fix |
|---|---|---|
| Disallow: / ships to production | Crawling of the whole site stops; organic visibility erodes | Remove the line immediately, resubmit your sitemap in Search Console, and request indexing for key pages. Recovery takes time as crawlers return — add robots.txt to your pre-deploy checklist so it never recurs |
| Blocking CSS/JS folders | Google renders broken pages and may misjudge content and mobile-friendliness | Remove the asset disallows and verify with Search Console’s URL Inspection that rendering looks right |
| Blocking a page to deindex it | URL stays in the index as a bare link, sometimes for a very long time | Unblock it, add a noindex tag, let Google recrawl and process the removal — then decide if blocking is even needed |
| Hiding sensitive URLs in robots.txt | The public file advertises exactly what you wanted hidden | Remove the lines and protect those areas with real authentication |
| Rules on the wrong subdomain | Your blog or shop subdomain crawls unmanaged (or stays blocked) | Give each subdomain its own correct robots.txt at its own root |
| Case mismatch in paths | Rules silently fail to match the real URLs | Copy exact paths from your live URLs into your rules, matching case precisely |
One encouraging note on recovery: robots.txt wounds are rarely permanent. Crawlers re-check the file regularly, so once you fix it, recovery begins on its own — faster if you resubmit sitemaps and nudge key pages through URL Inspection. The scar tissue becomes a checklist, and the checklist means it never happens twice.
Where does robots.txt fit in your bigger marketing picture?
Here’s some perspective before you go: robots.txt is plumbing. Essential, occasionally dramatic when it bursts, but not the thing your audience ever sees. You get the file right once, re-check it on deploys, and then your energy belongs to the work that actually grows an audience — publishing content worth finding and putting it in front of people consistently.
That second half is where I’ll be honest about what we build: SocialBlaze is an organic social media management platform, not an SEO tool — it won’t edit your robots.txt or crawl your site. What it does is make sure every article you publish actually reaches people: schedule and auto-publish across Instagram, Facebook, LinkedIn, Pinterest, YouTube, Threads, Bluesky, Mastodon, and more, then watch engagement from one dashboard. Search brings readers who are looking; social brings readers who didn’t know to look. Healthy sites run both.
Get your content seen beyond the search results
You’ve tidied the plumbing — now let people find the house. SocialBlaze schedules and auto-publishes your content across every major social network from one calendar, with analytics and a unified inbox built in — all on the Free Forever plan.
FAQ: how to use robots txt
Where does the robots.txt file go?
At the root of your domain: yoursite.com/robots.txt, exactly there and nowhere else. Each subdomain needs its own file at its own root — the main domain’s file doesn’t cover blog.yoursite.com or shop.yoursite.com. You can check any site’s file (including your own) by typing the URL straight into a browser.
Does robots.txt stop a page from appearing in Google?
Not reliably. Robots.txt blocks crawling, but a blocked URL can still be indexed if other pages link to it — it may show up in results as a bare link with no description. To actually keep a page out of search results, leave it crawlable and add a noindex meta tag, which Google honors once it recrawls the page.
Can I use robots.txt to hide private or sensitive pages?
No — and trying makes things worse. The file is publicly readable by anyone, and malicious bots ignore its rules entirely, so listing a sensitive path simply advertises it. Protect private areas with real access control: logins, password protection, or IP restrictions. Robots.txt is for crawl management of harmless pages, never for secrecy.
Will robots.txt changes improve my rankings?
Not directly, for most sites. A correct file prevents damage (like accidentally blocking your whole site or your CSS and JavaScript), and on very large sites it helps crawlers spend attention wisely. But for a typical small or mid-size site, robots.txt tuning isn’t a growth lever — clean file, deploy checklist, then invest your time in content and promotion.
Should I block ChatGPT and other AI crawlers in robots.txt?
It’s a business choice rather than an SEO tactic. Blocking AI training crawlers limits how your content feeds AI models; blocking AI search crawlers may reduce your visibility in AI-generated answers. Reputable AI companies’ crawlers generally respect robots.txt, but verify the current user-agent names in each company’s documentation before writing rules — the roster changes often.
Frequently Asked Questions
Social Blaze provides a comprehensive suite of features including social media scheduling, analytics, content libraries, team collaboration tools, RSS feed automation, and a browser extension to streamline your social media strategy.
Absolutely! Social Blaze is designed to cater to both small businesses and larger agencies, offering customizable solutions to fit various needs, whether you’re managing a single account or multiple clients.
Our AI assistant takes the hassle out of content creation by creating AI post content for you, think of it as your social media sidekick, saving you time while helping you level up your strategy with smart insights.
Yes! Social Blaze offers various integrations with popular platforms and tools, allowing you to streamline your workflow and enhance your social media management experience seamlessly.