Robots.txt: A Complete SEO Configuration Guide
Robots.txt is a plain text file that sits at the root of your domain and tells search engine crawlers which parts of your site they are and are not allowed to access. It follows the Robots Exclusion Protocol — a standard that every major search engine respects. When configured correctly, it steers crawlers efficiently through your site. When misconfigured, it can silently block pages you need indexed, or leave crawlers wasting budget on pages that add no value.
⠀
What Robots.txt Is — and What It Is Not
⠀
Robots.txt controls crawl access, not indexing. This is the most misunderstood aspect of the file. Blocking a URL in robots.txt prevents Googlebot from fetching its content, but it does not prevent the URL from appearing in search results. If another site links to a blocked URL, Google can still index that URL based on the link — it just cannot see the page's content.
Robots.txt is also not a security tool. Do not use it to hide sensitive content. Any determined user can read your robots.txt file directly in a browser. For content you genuinely need to restrict from public access, use authentication, server-level restrictions, or a noindex directive combined with proper access controls.
⠀
Where to Find Your Robots.txt File
⠀
Your robots.txt file is always located at the root of your domain: https://yourdomain.com/robots.txt. There is no other valid location. A file placed in a subdirectory (/blog/robots.txt) is not recognized as a valid robots.txt by search engines — it will simply be treated as a regular page.
To check whether your site has a robots.txt file, type your domain followed by /robots.txt into your browser's address bar. If you see a plain text file, the file exists and is reachable. If you get a 404, your site does not have one — which is fine, as the absence of a robots.txt file is interpreted as full crawl permission.
⠀
Understanding Robots.txt Directives
⠀
User-agent
⠀
The User-agent directive specifies which crawler a rule applies to. User-agent: * applies to all crawlers. User-agent: Googlebot applies only to Google's crawler. You can write separate rule blocks for different crawlers in the same file.
User-agent: * Disallow: /admin/ User-agent: Googlebot Disallow: /staging/
⠀
Disallow
⠀
The Disallow directive tells a crawler not to access a specific path. An empty Disallow value means no restrictions apply. A / disallows everything. Paths are matched from the beginning of the URL, so Disallow: /blog blocks /blog, /blog/post-1, /blog/category/, and any other URL starting with /blog.
Disallow: / ← blocks the entire site Disallow: /admin/ ← blocks only /admin/ and its subdirectories Disallow: ← no restriction (allows everything)
⠀
Allow
⠀
The Allow directive overrides a Disallow rule for a specific path within an already-blocked directory. It is most useful when you want to block a directory but allow specific files within it.
User-agent: * Disallow: /private/ Allow: /private/public-document.pdf
⠀
Sitemap Declaration
⠀
You can and should declare your sitemap URL in robots.txt. This is a separate directive that helps crawlers find your sitemap without relying solely on Google Search Console submission.
Sitemap: https://yourdomain.com/sitemap_index.xml
⠀
Multiple Sitemap lines are valid if you have more than one sitemap file.
⠀
What to Block with Robots.txt
⠀
Robots.txt is most useful for preventing crawlers from wasting time on pages that provide no indexing value. The most impactful categories to block are:
Admin and backend pages — /wp-admin/, /admin/, /dashboard/ — these should be blocked and are typically not publicly accessible anyway
Internal search result pages — pages like /?s=query or /search?q=query generate near-infinite duplicates and consume crawl budget without adding indexable value
URL parameter variations — if your site generates URLs with tracking parameters, sorting parameters, or filters that create duplicate content (?sort=price&color=red), block these to prevent crawl budget waste
Staging and development environments — if your staging environment is accessible at a subdomain, block it entirely to prevent accidental indexing of unfinished content
Duplicate content paths — printer-friendly versions, RSS feed paths that duplicate page content, and similar technical duplicates
⠀
At Blakfy, we use Screaming Frog to crawl client sites before auditing their robots.txt rules. Comparing what is blocked against what is actually being indexed reveals the gaps that cost crawl budget and dilute site quality signals.
⠀
Common Robots.txt Mistakes
⠀
Blocking CSS and JavaScript Files
⠀
Google renders pages like a browser. If your CSS and JavaScript files are blocked in robots.txt, Googlebot cannot see how your pages actually look — it gets raw, unstyled HTML. This can cause Google to misinterpret your page structure, miss content loaded by JavaScript, and fail to properly evaluate your site's mobile-friendliness. Never block your CSS, JS, or font files.
Blocking the Entire Site
⠀
A single line — Disallow: / under User-agent: * — blocks every crawler from every page on your site. This appears in new WordPress installations before a site launches (it maps to the "discourage search engines" setting in WordPress). Forgetting to remove it after launch is one of the most common causes of sites failing to appear in search results after going live.
Using Noindex and Disallow Together
⠀
If you block a URL with Disallow and also add a noindex directive to that page, the noindex is effectively invisible. Googlebot cannot read the noindex tag because it cannot access the page. If you want Google to recognize a noindex directive, the page must be crawlable. Use noindex only on pages Googlebot can reach.
Blocking Paginated URLs Incorrectly
⠀
Pagination is often handled poorly in robots.txt. Blocking all paginated URLs (Disallow: /?page=) can prevent Google from discovering products or posts that only appear on page 2 and beyond. Be deliberate about what you block — pagination itself is not a problem worth blocking unless the pages are genuinely duplicate content.
⠀
How to Test Your Robots.txt in Google Search Console
⠀
Open Google Search Console and select your property.
Navigate to Settings > robots.txt (in the Legacy tools section, or access directly via the URL inspection workflow).
Alternatively, use the standalone robots.txt Tester available in GSC.
Enter the URL you want to test in the tester field.
Click Test — the tool shows whether that URL is blocked or allowed for Googlebot specifically.
⠀
The tester highlights which rule in your file is being applied to the URL, which makes it easier to debug unexpected blocks. Note that the tester reflects the robots.txt file Google has cached, not necessarily the current version — click "Submit to Google" after making changes to prompt a refetch.
You can also test robots.txt behavior using Screaming Frog by configuring it to respect the robots.txt file and comparing the crawl results against a crawl without restrictions.
⠀
Robots.txt vs. Meta Robots vs. X-Robots-Tag
⠀
These three tools are often confused, but they serve distinct functions:
Robots.txt — controls whether a URL can be crawled at all
Meta robots tag (<meta name="robots" content="noindex">) — placed in the HTML <head>, controls indexing of a crawlable page
X-Robots-Tag — an HTTP response header that controls indexing and can be applied to non-HTML files like PDFs
⠀
The correct approach depends on what you want to achieve. To keep a page out of search results while allowing Google to crawl it (so it can read the noindex instruction), use a meta robots tag or X-Robots-Tag — not robots.txt. To prevent crawlers from accessing a path entirely (saving crawl budget), use robots.txt — but accept that the URL can still be indexed via external links.
⠀
⠀
FAQ
⠀
Does blocking a URL in robots.txt remove it from Google's index?
No. Disallowing a URL prevents Googlebot from crawling it, but if that URL has external links pointing to it, Google can still index it based on those signals. To remove a URL from the index, use a noindex directive on a crawlable page, or use Google Search Console's URL removal tool for temporary suppression.
Can I use wildcards in robots.txt?
Yes, but with limitations. The * wildcard in a path matches any sequence of characters, and $ matches the end of a URL. For example, Disallow: /*?* blocks any URL containing a ? character (useful for blocking all parameterized URLs). Google supports these patterns; not all crawlers do.
How quickly does Google pick up robots.txt changes?
Google typically recrawls your robots.txt file every 24 hours, but it can take longer depending on your site's crawl frequency. After making changes, use Google Search Console to submit the updated file for faster processing. The change will not be instant — allow 24-48 hours for the update to propagate.
Should every website have a robots.txt file?
Not necessarily. A missing robots.txt is interpreted as full crawl permission, which is the default state most sites want. You only need a robots.txt file if you have specific pages or directories you want to restrict from crawlers. That said, declaring your sitemap via robots.txt is a good practice regardless.



