top of page

Robots.txt: How to Configure It Without Blocking Google

Jan 13
4 min read

The robots.txt file is a plain text file placed at the root of your website (yourdomain.com/robots.txt) that instructs search engine crawlers which pages or directories they are permitted to access. It is one of the simplest files on a website and one of the most dangerous to misconfigure — a single incorrect Disallow rule can block Google from crawling your entire site.

Understanding robots.txt correctly means understanding what it does and doesn't do: it controls crawl access, not indexation. A page blocked in robots.txt won't be crawled, but if other sites link to it, Google may still index it based on those external links (without being able to read the content).

⠀

How Robots.txt Works

⠀

The robots.txt protocol is based on the Robots Exclusion Standard. Each set of rules in the file has two components:

User-agent: Specifies which crawler the rules apply to. User-agent: * applies to all crawlers. User-agent: Googlebot applies only to Google's crawler.

Disallow / Allow directives: Tell crawlers what they can and cannot access.

User-agent: * Disallow: /admin/ Disallow: /thank-you/ Allow: / Sitemap: https://yourdomain.com/sitemap.xml

⠀

This example blocks all crawlers from /admin/ and /thank-you/, allows everything else, and points to the sitemap.

Important: robots.txt is a request, not a command. Reputable crawlers (Google, Bing) respect it. Malicious bots may not.

⠀

What to Block in Robots.txt

⠀

⠀

⠀

A well-configured robots.txt blocks access to pages that should never appear in search results and pages that waste crawl budget without providing indexable value:

Admin and backend areas:

Disallow: /admin/ Disallow: /wp-admin/ Disallow: /login/ Disallow: /dashboard/

⠀

These pages should never appear in search results. Blocking them prevents crawlers from wasting time on pages with no SEO value.

Thank-you and confirmation pages:

Disallow: /thank-you/ Disallow: /order-confirmation/ Disallow: /checkout/complete/

⠀

Post-conversion pages have no SEO value and shouldn't be indexed.

Internal search result pages:

Disallow: /search/ Disallow: /?s=

⠀

Internal search results create near-infinite duplicate thin-content pages. Blocking these prevents crawl budget waste.

Staging and development URLs (if on the same domain):

Disallow: /staging/ Disallow: /dev/

⠀

What NOT to block:

  • Service pages, product pages, blog posts — these need to be crawled and indexed

  • CSS and JavaScript files — blocking these prevents Google from rendering your pages correctly, which can hurt rankings

  • Pages you've blocked with noindex — you don't need to block them in robots.txt if you've already added noindex; it's redundant and can actually prevent Google from seeing your noindex directive

⠀

⠀

Common Robots.txt Mistakes

⠀

Accidentally blocking everything:

Disallow: /

⠀

This blocks all crawlers from all pages. It's the most catastrophic robots.txt error. If you see this in your file, fix it immediately.

Blocking CSS and JavaScript:

Old robots.txt recommendations sometimes advised blocking /wp-content/ or asset directories. This prevents Google from rendering your pages and accurately assessing their content and design. Allow all CSS, JavaScript, and image assets.

Blocking pages you want indexed:

Any page you want in Google's index must not be blocked in robots.txt. Verify this by checking the URL Inspection tool in Google Search Console for pages you suspect may be blocked.

Contradicting noindex with robots.txt blocking:

If you block a page in robots.txt, Google can't crawl it — which means Google can't see the noindex tag on the page. If you want a page definitively not indexed, either: (a) block it in robots.txt (Google will typically de-index it over time), or (b) allow crawling but add a noindex meta tag. Doing both is contradictory.

⠀

Testing Your Robots.txt

⠀

⠀

⠀

Google Search Console robots.txt Tester:

In Search Console, the robots.txt tester (Settings → robots.txt) allows you to test any URL against your current robots.txt rules. Enter a URL and it shows whether Googlebot can access it — and which specific rule is blocking or allowing it.

Manual inspection:

Access yourdomain.com/robots.txt in your browser. Verify the file is accessible (returns 200) and contains the correct rules.

Screaming Frog:

Screaming Frog reads your robots.txt during crawls and marks blocked URLs. Crawl your site with Screaming Frog and filter by "Blocked by Robots.txt" to see what you're preventing Google from accessing.

Google Search Console Coverage report:

URLs blocked by robots.txt appear in the Coverage report under "Excluded" → "Blocked by robots.txt". If important pages appear here, your robots.txt is blocking pages that should be crawled.

⠀

Robots.txt and the Sitemap

⠀

Best practice is to reference your XML sitemap in robots.txt:

Sitemap: https://yourdomain.com/sitemap.xml

⠀

This allows any crawler that reads robots.txt to discover your sitemap automatically. Place this line at the bottom of the file after your Disallow rules.

Blakfy configures and audits robots.txt files for clients — ensuring crawlers have access to all pages that should be indexed while blocking the administrative and low-value sections that waste crawl budget.

⠀

⠀

Frequently Asked Questions

⠀

Does robots.txt affect Google rankings?

Robots.txt doesn't directly affect rankings — but a misconfigured robots.txt that blocks important pages from crawling will prevent those pages from ranking. The indirect impact of an incorrect robots.txt can be severe: pages that can't be crawled can't be indexed, and pages that aren't indexed can't rank. Verifying robots.txt correctness is a foundational technical SEO check.

Can I block specific bots but allow Googlebot?

Yes. Use User-agent directives to apply rules selectively:

User-agent: AhrefsBot Disallow: / User-agent: Googlebot Allow: /

⠀

This blocks Ahrefs' crawler while allowing Googlebot. Note that malicious scrapers typically ignore robots.txt regardless.

My site was ranking, then traffic dropped — could robots.txt be the cause?

Yes, and this is one of the first things to check after an unexpected traffic drop. A CMS update, site migration, or developer change can accidentally modify robots.txt to block crawling. Check your robots.txt immediately if rankings drop sharply after a site change, particularly if the traffic drop affects the entire site rather than individual pages.

Should I block AI crawlers in robots.txt?

Some site owners choose to block AI training crawlers (GPTBot, Claude-Web, Common Crawl) in robots.txt to prevent their content from being used for AI training. This is a policy decision, not an SEO decision. Blocking these crawlers has no impact on Google search rankings. If content scraping for AI training is a concern, add the relevant User-agent blocks with Disallow: /.

bottom of page