Skip to content
MarketHQ
Glossary

Crawl Budget

Quick answer

Crawl budget is the limited number of pages on a site that a crawler, whether a classic search bot or an AI company's bot, will actually fetch within a given stretch of time. A large or poorly organized site can have pages that never get crawled at all, simply because the budget ran out first.

Why it matters

A page that never gets crawled can't be cited by an AI engine or ranked by a search engine, no matter how good its content is, so crawl budget acts as a hard ceiling under everything else a content or SEO team does. For a site with thousands of pages and a lot of low-value duplication, a chunk of that limited budget gets wasted on pages nobody needed crawled in the first place, at the expense of the pages that actually matter. Crawl budget is mostly a concern for large or fast-changing sites, but the underlying idea applies to any site: crawlers have limited time, and pages that are slow, duplicated or buried deep in the structure are the first ones to be skipped.

How to measure it

Check server logs or a crawl-analytics tool for which bots are visiting, how often, and which pages they're actually fetching versus skipping. An accurate sitemap and a robots.txt file that blocks low-value pages, such as admin routes or duplicate filters, both help steer a limited crawl budget toward the pages that matter most to AI and search visibility alike. Also look at how quickly new or updated pages are first fetched after publishing. A long delay for important pages, alongside heavy activity on parameter URLs, redirects or error pages, shows where the budget is leaking and what to consolidate or block.

Example

A site with thousands of auto-generated filter pages notices its newest, most important product page hasn't been crawled in weeks. Checking server logs shows crawlers are spending most of their budget on filter combinations nobody searches for, a problem fixed by blocking those low-value pages in robots.txt so the budget reaches the pages that actually matter, and the new page gets crawled within days instead of weeks.

Frequently asked questions

Does every site need to worry about crawl budget?

Mostly larger sites with thousands of pages. A small site with a few dozen pages rarely runs into a real crawl-budget limit and can usually treat this as a non-issue.

Is crawl budget the same for AI bots and search bots?

Each bot, such as GPTBot, PerplexityBot or Googlebot, sets its own crawl budget independently, so a page favored by one crawler isn't guaranteed to be favored by another.

Can robots.txt help manage crawl budget?

Yes. Blocking low-value or duplicate pages in robots.txt frees up budget for a crawler to spend on the pages that actually matter.

How can a team tell if crawl budget is actually a problem?

Server logs showing important pages going uncrawled for an unusually long stretch, alongside a lot of crawler activity on low-value pages, is the clearest sign the budget is being spent in the wrong place.

Related terms

Related

MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.