top of page

Screaming Frog SEO Spider: Complete Beginner to Advanced Guide

Oct 12, 2028
5 min read

Screaming Frog SEO Spider is the industry-standard technical SEO crawling tool used by professional SEOs worldwide. It crawls websites like a search engine bot, collecting data on every URL, and surfaces a comprehensive range of technical issues in a tabular interface. Whether you're doing a quick 500-page crawl on the free version or a full enterprise crawl of millions of URLs, Screaming Frog remains the most detailed technical SEO analysis tool available.

⠀

What Screaming Frog Does: Screaming Frog Seo

⠀

Screaming Frog is a desktop application (Windows, Mac, Ubuntu) that sends HTTP requests to your website, follows links, and collects data from every URL it finds. Unlike cloud-based tools, it runs locally on your machine, using your IP and your internet connection.

What it crawls and collects:

  • All HTML pages, images, JavaScript files, CSS files, and other resources

  • HTTP status codes (200, 301, 302, 404, 500, etc.)

  • Title tags and meta descriptions

  • Heading structure (H1-H6)

  • Canonical tags

  • Noindex and nofollow directives

  • Response times

  • Word counts

  • hreflang attributes

  • Structured data (schema markup)

  • Core Web Vitals data (via Google's API integration)

  • Internal and external links from each page

⠀

Free vs Paid:

The free version crawls up to 500 URLs per site — sufficient for small sites. The paid license (~$259/year) removes the URL limit and adds advanced features including JavaScript rendering, Google Analytics integration, Google Search Console integration, and custom extraction.

⠀

Basic Setup and Your First Crawl ve Screaming Frog Seo

⠀

⠀

⠀

Starting a crawl:

  1. Open Screaming Frog

  2. Enter the domain you want to crawl in the address bar (e.g., https://example.com)

  3. Click Start

⠀

By default, Screaming Frog crawls all internal URLs it discovers starting from the given URL, following links as it goes.

Recommended initial settings:

Before crawling, configure a few defaults:

Configuration > Spider > Crawl > Check these boxes:

  • "Check Canonicals" — crawls canonical URLs even if they're not linked

  • "Check hreflang" — validates hreflang configurations

  • "Crawl Outside of Start Folder" — prevents crawls from being restricted to a subfolder

⠀

Configuration > User-Agent > Select "Googlebot (Desktop)" to simulate Google's crawl.

Configuration > Speed > Set crawl speed based on your server's capacity. 1-3 requests/second is safe for most sites. Higher speeds can overload small servers.

Excluding URLs:

Configuration > Exclude — add URL patterns you want to exclude (e.g., /cart/, /checkout/, /wp-admin/). This prevents crawling non-indexable or administrative URLs.

⠀

The Most Important Reports and How to Use Them

⠀

Response Codes tab:

The most critical starting point. Filter by status code:

  • 200 OK: Your successfully crawling pages

  • 301 Redirects: All redirected pages — check for redirect chains

  • 404 Not Found: Broken pages — especially important if they have internal links or backlinks

  • 500 Server Errors: Server-side issues that need immediate attention

⠀

Page Titles tab:

Filter for:

  • Missing (empty title tags)

  • Duplicate (same title on multiple pages)

  • Too long (over 60 characters)

  • Too short (under 30 characters)

  • Multiple (pages with more than one title tag)

⠀

Meta Descriptions tab:

Same filters as Page Titles: Missing, Duplicate, Too Long, Too Short.

H1 tab:

Filter for: Missing, Duplicate, Multiple (pages with more than one H1).

Canonicals tab:

Filter for:

  • Pages with no canonical tag

  • Canonical URLs pointing to a different URL (non-self-referential)

  • Canonical chains (A → B → C)

  • Canonicals pointing to redirects or 404 pages

⠀

Images tab:

Filter for missing alt text. Click "Export" to get a list of all images without alt attributes.

Internal Links tab:

Shows the number of inlinks (links pointing to each page from within the site). Sort by "Inlinks" ascending to find orphan pages (pages with zero or very few internal links).

⠀

Advanced Screaming Frog Configurations

⠀

JavaScript Rendering:

By default, Screaming Frog crawls raw HTML without JavaScript execution. For JavaScript-heavy sites (React, Angular, Vue), enable JavaScript rendering:

Configuration > Spider > Rendering > JavaScript

This uses a headless Chrome instance to render pages with JavaScript before collecting data. JavaScript rendering is slower and more resource-intensive but essential for accurate crawls of modern JS frameworks.

Custom Extraction:

Configuration > Custom > Extraction — extract any data from page HTML using CSS selectors, XPath, or regex. Uses:

  • Extract structured data values (e.g., prices, ratings)

  • Extract custom meta tags used by your CMS

  • Extract specific content elements for audit purposes

⠀

Google Analytics Integration:

Configuration > API Access > Google Analytics — connect to GA4 to add organic sessions data to your crawl. This allows you to filter and sort pages by actual traffic, making it easy to prioritize high-traffic pages for optimization.

Google Search Console Integration:

Configuration > API Access > Google Search Console — add GSC data including clicks, impressions, CTR, and position to your crawl data. Combined with crawler data, this is extremely powerful for identifying high-impression pages with technical issues.

⠀

Analyzing Crawl Data and Exporting

⠀

⠀

⠀

Key analysis workflows:

Finding orphan pages:

Go to Internal > All Internal URLs. Filter by "Inlinks: 0." These pages have no internal links pointing to them — Google may not discover or adequately crawl them.

Finding redirect chains:

Response Codes > Redirects. Sort by "Redirect Chain" column. Chains longer than 1 hop (A→B→C) should be simplified to direct redirects.

Identifying thin content:

Crawl > Filter > HTML pages. Sort by "Word Count" ascending. Pages with very low word counts (under 200) are thin content candidates.

Duplicate content detection:

Crawl > Duplication tab. Screaming Frog uses hash-based comparison to identify near-duplicate pages.

Exporting data:

Every tab in Screaming Frog can be exported to CSV via the "Export" button. For large audits, export all relevant tabs and analyze in Excel or Google Sheets. The Bulk Export option (right-click) allows exporting multiple tabs at once.

For client reports, Screaming Frog's built-in Crawl Report (Reports > Crawl Overview) generates a PDF summary of key metrics.

Blakfy uses Screaming Frog as the primary technical crawling tool in all SEO audits, integrating crawl data with GSC and GA4 for comprehensive technical analysis.

⠀

Frequently Asked Questions

⠀

How long does a Screaming Frog crawl take?

Crawl time depends on the number of pages and your crawl speed setting. A 1,000-page site at 3 requests/second takes roughly 5-10 minutes. A 100,000-page site may take several hours. For very large sites (1M+ pages), configure list mode to crawl only specific URLs (from your sitemap) rather than a full spider crawl.

Can Screaming Frog crawl password-protected pages?

Yes. Go to Configuration > Authentication to set HTTP authentication credentials for password-protected sections. For CMS login walls (like WordPress admin), you can use a cookie-based authentication approach via Configuration > Cookie.

Is the free version of Screaming Frog worth using?

Absolutely. The free 500 URL limit makes it practical for small business sites, blog audits, and learning the tool. For sites with more pages or for professional auditing work, the paid license at ~$259/year is one of the best value investments in SEO tooling — cheaper than one month of Ahrefs or SEMrush, and more powerful for technical crawling than either platform.

bottom of page