SEO

Crawling

Also known as: Web crawling, Crawl, Spidering

Crawling is the process in which search engine bots (Googlebot, Bingbot and so on) automatically retrieve web pages, parse their HTML, extract new links and place those links in a queue for further crawling. Crawling is the first stage in the pipeline: crawl → render → index → rank. Anything that is not crawled does not appear in the index and cannot rank. On large websites, crawl budget — the crawl capacity Google assigns to a domain — is a limiting factor.

Understanding crawl budget

Google allocates each domain a crawl budget, which is made up of two components: the crawl rate limit (how many requests per second can the server handle?) and crawl demand (how important does Google consider the domain and its content?). On small sites (< 10,000 URLs) budget is not an issue — Google crawls everything. From around 100,000 URLs onwards budget becomes relevant: not all URLs are crawled, and some only every few weeks. On very large sites (millions of URLs) budget management is essential.

What consumes crawl budget

Controlling crawling

Example from practice

Example: A shop with 38,000 URLs had the following crawl statistics in GSC: Googlebot crawled 22,000 URLs a day, 60% of which were filter and sorting parameter variants. Setting canonical tags on all parameter URLs, plus Disallow: /search? and targeted wildcard blocks in robots.txt, reduced the crawl to 8,500 URLs a day, all of them canonical products. New products now appear in the index after an average of 4 days instead of 14 — with the same crawl effort from the bot.

Frequently asked questions

What is crawling in SEO?
Crawling is the process in which search engine bots (Googlebot, Bingbot) visit websites and read their content. Crawling is the prerequisite for indexing — a page that has not been crawled cannot appear in search results.
What is a crawl budget?
The crawl budget is the number of URLs Googlebot crawls on a site within a given period. Google balances server load against the need for updates. Small sites under 1,000 URLs are almost never budget-limited, whereas large sites with 100,000+ URLs have to prioritise.
How can I see crawl behaviour?
In Google Search Console under "Settings → Crawl stats". The report shows the number of crawl requests per day, response times and error status codes. An alternative is to filter server logs for the user agent "Googlebot".
How do you control crawling?
Through four levers: (1) robots.txt to exclude sections, (2) noindex to exclude individual pages, (3) an XML sitemap for prioritisation, (4) internal linking for signal strength per URL. The crawl rate limit setting in GSC is deprecated (2024).

Used in Rankmio for

Crawl statistics and crawl budget audit

Go to the feature →

Last updated: 2026-06-17  ·  Browse all glossary entries

Free SEO & GEO Check

SEO score, AI visibility and citability of your website in 30 seconds — no registration required.

Check for free now

Ready to optimize your website?

Register for free, get 10 credits and start right away.

Register now