Also known as: Web crawling, Crawl, Spidering
Crawling is the process in which search engine bots (Googlebot, Bingbot and so on) automatically retrieve web pages, parse their HTML, extract new links and place those links in a queue for further crawling. Crawling is the first stage in the pipeline: crawl → render → index → rank. Anything that is not crawled does not appear in the index and cannot rank. On large websites, crawl budget — the crawl capacity Google assigns to a domain — is a limiting factor.
Google allocates each domain a crawl budget, which is made up of two components: the crawl rate limit (how many requests per second can the server handle?) and crawl demand (how important does Google consider the domain and its content?). On small sites (< 10,000 URLs) budget is not an issue — Google crawls everything. From around 100,000 URLs onwards budget becomes relevant: not all URLs are crawled, and some only every few weeks. On very large sites (millions of URLs) budget management is essential.
Example: A shop with 38,000 URLs had the following crawl statistics in GSC: Googlebot crawled 22,000 URLs a day, 60% of which were filter and sorting parameter variants. Setting canonical tags on all parameter URLs, plus Disallow: /search? and targeted wildcard blocks in robots.txt, reduced the crawl to 8,500 URLs a day, all of them canonical products. New products now appear in the index after an average of 4 days instead of 14 — with the same crawl effort from the bot.
Crawl statistics and crawl budget audit
Free SEO & GEO Check
SEO score, AI visibility and citability of your website in 30 seconds — no registration required.
Register for free, get 10 credits and start right away.
Register now