What is Index Bloat? — Whiteboard Friday
by ronyau
Deep dive into index bloat - a critical SEO challenge affecting medium to large websites. Learn how to identify URLs consuming your index quota without delivering traffic, understand the difference between crawl budget and index bloat, and discover practical solutions for cleanup. This Whiteboard Friday video helps you assess your site's index health and implement effective remediation steps, f...
Deep dive into index bloat - a critical SEO challenge affecting medium to large websites. Learn how to identify URLs consuming your index quota without delivering traffic, understand the difference between crawl budget and index bloat, and discover practical solutions for cleanup. This Whiteboard Friday video helps you assess your site's index health and implement effective remediation steps, from content consolidation to proper URL handling.
Click on the whiteboard image above to open a high-resolution version!
Happy Friday, Moz fans. Today I want to talk about index bloat.
So this is a pretty common problem affecting especially large, but also sometimes medium-sized sites. And I'd say this is definitely something that you should have looked into if you work for a medium or a larger site. It's definitely something you should have looked into at least once. It does affect a lot of sites. It's worth checking whether this might affect you. This is something that I and a lot of other SEOs have seen very good results with both for a long time and very recently. And despite that, it's something that I think is relatively poorly codified and talked about in the industry, there are some reasons for that which I'll come on to in a moment.
Understanding index bloat
But before we get into all of that, I just want to explain this. So I put this diagram in just to give a bit of context, so that what I say next makes sense. So this outer box, this diagram as a whole is all of the URLs on your site, all of the URLs that might exist, that could exist, including parameters that no one has tried before, this kind of thing, the maximum possible set of URLs that would return a 200 response code and a valid page.
And then I've got smaller sets within that, sort of subsets of URLs. So the next one down is Google discovered URLs. So if Google has seen the URL – they might have not crawled it, they might have not indexed it, but they've seen the URL, they know it exists. That's sort of your next step down. And if you've got a big difference between the red box and the blue box, that probably indicates some kind of crawl budget problem. But that's not what we're talking about today.
If you've got a URL that's discovered, it might not necessarily be indexed. So indexed URLs is another smaller set. If you've got a URL that's discovered but not indexed, again there might be some reasons for that. Google might suspect that the page is unimportant based on other signals. You might have said not to index it. You might have shown them a noindex tag or something like this. So that, again, is a smaller set. Again, we're not necessarily speaking about that gap today.
And then you've got indexed versus pages with non-trivial traffic. Now what counts as non-trivial traffic to a page might vary from site to site. You might have your own idea of this. But a big gap between the number of URLs that are indexed and the number of URLs that are getting any kind of meaningful non-zero traffic, if that's a big gap, that would suggest an index bloat problem, and that's what I want to talk about today.
Never miss an issue impacting traffic on your site
Find and fix technical SEO issues fast with Moz Pro.
What index bloat is not
So before I get into that, just to make it totally clear, I want to quickly disambiguate a couple of things I mentioned there. So I'm not talking about crawl budget. As I mentioned before, that's when you've got a lot of URLs that Google just isn't going to crawl at all. Perhaps you're producing them too quickly. You've got a very large number, a huge number of URLs on your site. This might affect news websites, for example, sometimes large forums.
I'm also not talking about cannibalization. Now that is a related concept. Often when you've got a huge number of indexed pages that aren't getting traffic, it's because their topics are too similar. But you could theoretically have a cannibalization problem on a site with three pages, if they're all about roughly the same thing. That's not really what I'm talking about today. I'm talking about a larger-scale problem.
So we're specifically talking about the difference between the yellow and green boxes I talked about earlier. So how many of all the URLs that are indexed, are there a large number that Google is not really bothering to send any meaningful traffic to or show up in search results?