What Is Indexing?

Indexing is the second stage of how search engines process web content, following crawling. Once Googlebot has fetched a page, Google's indexing systems analyse the page's content, images, video, structured data, and signals such as canonical tags and noindex directives. If the page meets Google's quality and technical requirements, it is added to the search index, a vast database that serves as the source for all search results.

Indexing is distinct from ranking. A page that is indexed has been stored in Google's database, but its ranking position in the SERP depends on many additional signals including relevance, authority, page experience, and content quality. You can be indexed for a query but still appear on page ten because other pages are deemed more relevant or authoritative.

Several factors can prevent a crawled page from being indexed. A noindex meta robots tag or HTTP header tells Google not to index the page. Duplicate content may cause Google to index the canonical version rather than the duplicate. Thin content with very little unique information may be excluded from the index if Google does not consider it useful to searchers. Soft 404 errors, where a page returns a 200 status but delivers a "not found" message, can also cause indexing to be withheld.

Google Search Console's Index Coverage report is the primary tool for monitoring indexing status across a website. It categorises pages as indexed, excluded, or having errors, and provides specific reasons for exclusion to help SEO practitioners diagnose and resolve coverage issues.

Indexing In Practice

The scenario below is an illustrative example, not a Juicy Designs client result. The figures indicate the scale of effect that indexing work typically produces, so treat them as indicative rather than measured.

Picture a Durban property agency that launches a new development listing page and notices it does not appear in Google search results after around two weeks. Using Search Console's URL Inspection tool, the agency might find that the page has not yet been indexed, with the reason given as "Discovered, currently not indexed," meaning Google found the URL but has not yet processed it fully, often because the page is new and Google has not yet allocated rendering resources to it.

To accelerate indexing, the agency could use the "Request Indexing" function in the URL Inspection tool, add the URL to its XML sitemap, and build internal links from the agency's homepage and other high-authority pages to the new listing. For frequently updated sites like property portals, ensuring a clean sitemap with accurate lastmod dates helps Google prioritise fresh content in its crawl and index queue.

South African businesses operating multilingual sites or e-commerce stores with faceted navigation frequently encounter index bloat, where thousands of low-value filter URLs get indexed. Applying noindex tags to faceted navigation pages and consolidating duplicate content with canonical tags helps keep the index clean and ensures crawl budget is allocated to pages that deliver ranking and traffic value.

Crawling versus indexing

Crawling and indexing are two separate steps, and confusing them causes wasted effort. Crawling is discovery: a search engine's bot fetches a page's content. Indexing is inclusion: after crawling, the engine analyses the page and, if it judges it worthy, stores it in the index so it can appear in results. A page can be crawled but not indexed, which is common and usually means the content was judged thin, duplicate or low value. Only indexed pages can rank, and only indexed pages are eligible for AI features such as Overviews. So being crawled is necessary but not sufficient; the goal is to be indexed.

Why pages do not get indexed

When Google crawls a page but declines to index it, the usual reasons are quality and duplication: the content is thin, near-identical to other pages, or offers little a searcher could not get elsewhere. Technical causes also block indexing, a noindex tag, a canonical pointing elsewhere, or the page being orphaned with no internal links. For large or new sites, Google may also simply be slow to get to lower-priority pages. The remedy depends on the cause: strengthen thin content, consolidate duplicates, remove accidental noindex or canonical errors, and add internal links so the page is clearly reachable and clearly worth including.

FAQ

Why is my page crawled but not indexed by Google?

Google may crawl a page without indexing it if it finds the content to be thin, duplicate, or of low quality. Other reasons include a noindex directive, a canonical tag pointing to a different URL, or the page being a near-duplicate of content already in the index.

How do I check if my pages are indexed in Google?

Use Google Search Console's Index Coverage report or the URL Inspection tool to check indexing status for specific pages. You can also search site:yourdomain.co.za in Google to get a rough count of indexed pages, though this is not exhaustive.

How do you get a page indexed faster?

Ensure the page is genuinely useful and free of noindex or canonical errors, link to it from other indexed pages, include it in your XML sitemap, and use Search Console's URL Inspection to request indexing. Quality and clear internal links matter most; requesting is only a nudge.

Does indexing guarantee ranking?

No. Indexing only makes a page eligible to appear in results; where it ranks then depends on relevance, quality and authority relative to competitors. A page must be indexed to rank, but many indexed pages rank poorly or not at all for competitive terms.

Want a team that knows these metrics cold?

Founder-led digital marketing for South African businesses since 2015. 4.9-star rated, 64+ clients, no long-term contracts.