What Is Indexability?

Indexability describes whether a page that has been crawled by a search engine qualifies for inclusion in its index. The index is the database from which search results are drawn. A page must be both crawlable and indexable to appear in search results. Crawlability handles access, while indexability handles eligibility.

Several factors determine whether a page is indexable. The most direct is the noindex directive, which can appear as a meta robots tag (<meta name="robots" content="noindex">) or as an HTTP response header (X-Robots-Tag: noindex). When either is present, Google will crawl the page but remove it from the index or decline to add it.

Canonical tags also affect indexability. A canonical tag tells Google that a given URL is a duplicate or variant of another URL. Google will typically index only the canonical URL and exclude all pages that point to it as their canonical. This is useful for managing duplicate content in e-commerce pagination or URL parameter variants, but misconfigured canonicals can accidentally exclude important pages from the index.

Content quality is a third indexability factor. Google evaluates whether a crawled page provides sufficient value to be worth including in the index. Pages with very thin content, those that are substantially similar to other indexed pages, or those with poor engagement signals may be excluded or de-prioritised. This is particularly relevant for South African SEO clients in competitive niches where Google's quality threshold is higher than expected.

Indexability In Practice

A Cape Town e-commerce store runs WooCommerce with faceted navigation, generating URLs like /products-colour=red&size=M&sort=price for every filter combination. These URLs have no unique content and their canonical tags were accidentally left pointing to themselves rather than to the clean base URL. The result: hundreds of parameter URLs appearing in the index, diluting the site's crawl budget and splitting ranking signals across near-identical pages. Fixing indexability here means setting correct self-referencing canonicals on base product pages and noindex on all filtered variants.

Another common South African scenario is the staging or development site that is launched to production with <meta name="robots" content="noindex,nofollow"> still in place. The site receives no organic traffic, Google Search Console's Index Coverage report shows zero indexed pages, and the team spends weeks investigating other causes before finding the single meta tag responsible.

To audit indexability, check the Index Coverage report in Google Search Console for pages listed under "Excluded." The most actionable exclusion reasons include "Excluded by 'noindex' tag," "Alternate page with proper canonical tag," and "Duplicate without user-selected canonical." Each tells you exactly why Google is withholding the page from the index and what needs to change.

What affects indexability

Indexability is whether a page can be included in a search engine's index once found. Several things prevent it. A noindex tag explicitly tells search engines not to list the page. A canonical tag pointing elsewhere signals that another URL is the version to index. Low quality or duplication can lead search engines to crawl a page but decline to index it, judging it not worth including. Being blocked from crawling in robots.txt can stop the engine seeing a noindex or the content at all. And a page with no internal links may not be discovered to be indexed in the first place. Indexability therefore depends on both technical directives and content quality, since a page must be permitted, discoverable and worth including.

How to improve indexability

Improving indexability means removing the barriers that keep worthy pages out of the index. Check for and remove accidental noindex tags and canonical tags pointing to the wrong URL, the most common technical culprits. Ensure important pages are not blocked in robots.txt, are linked internally so they are discoverable, and are included in the XML sitemap. Then address quality, since search engines decline to index thin or duplicate pages: strengthen weak content, consolidate near-duplicates, and give each page a clear purpose. Use Search Console to see which pages are indexed and why others are not, then act on the specific reason. The goal is that every page you want found is both technically permitted and genuinely worth indexing.

FAQ

What is the difference between crawlability and indexability?

Crawlability is about whether a bot can access and retrieve a page. Indexability is about whether that retrieved page qualifies for inclusion in the search index. A page can be perfectly crawlable but not indexable if it carries a noindex directive, is canonicalised to another URL, or contains content Google considers too thin or duplicate to be worth indexing.

Why would Google crawl a page but not index it?

Google crawls pages to evaluate them but decides not to index a page when it finds a noindex tag, a canonical tag pointing to a different URL, very thin content, significant duplicate content already indexed elsewhere, or low-quality signals such as no inbound links and slow load times. The Index Coverage report in Google Search Console shows the specific reason for any non-indexed URL.

Want a team that knows these metrics cold?

Founder-led digital marketing for South African businesses since 2015. 4.9-star rated, 64+ clients, no long-term contracts.