What Is Noindex?

Noindex is a technical SEO directive that tells search engines not to add a specific URL to their search index. It is most commonly applied using the robots meta tag placed in the HTML head of a page: <meta name="robots" content="noindex">. Alternatively, it can be set via an X-Robots-Tag in the HTTP response headers, which is useful for non-HTML files like PDFs.

When Googlebot or another crawler visits a page and reads a noindex directive, it will not include that page in search results. The page may still be crawled to read the tag, and it can still receive internal links and pass link equity, but it will not appear in search results for any query. This is distinct from using robots.txt, which blocks crawling entirely and prevents Google from reading the page at all.

Noindex is used strategically across many website types to improve the quality of indexed content. A South African e-commerce site might noindex checkout pages, thank-you pages, internal search results, user account dashboards, and paginated category pages beyond the first page. By removing these low-value pages from the index, the site concentrates its crawl budget on pages that actually have ranking potential.

Noindex can be combined with follow to allow Google to crawl links on the page even while keeping the page itself out of the index. The full tag would read: <meta name="robots" content="noindex, follow">. This is often appropriate for tag archives, filtered pages, and other generated URLs that aggregate content rather than presenting unique content of their own.

Noindex In Practice

The scenario below is an illustrative example, not a Juicy Designs client result. The details indicate the scale of effect that noindex work typically produces, so treat them as indicative rather than measured.

Picture a Johannesburg property portal that generates thousands of search results pages when users filter by area, price, and property type. URLs like /properties/-area=sandton&bedrooms=3&price=2m would be created dynamically and could be indexed by Google, producing many near-identical pages in the index and diluting the site's SEO performance for the core listing pages.

The fix would be to apply a noindex tag to all parameter-driven search results pages, while keeping the individual property listing pages indexed. Over time this would remove thousands of low-value pages from Google's index, allowing crawl budget to be allocated to individual property pages, neighbourhood guide articles, and the main category landing pages. After an implementation of this kind, a portal's primary category pages would typically see improved ranking consolidation as Google recognises them as the canonical source of content for each area.

Noindex versus robots.txt

These two controls are often confused and do opposite jobs. A noindex tag allows a page to be crawled but tells search engines not to list it in results. A robots.txt disallow tells crawlers not to fetch the page at all. The crucial interaction is that a page blocked in robots.txt cannot be crawled, so search engines never see its noindex tag, meaning the two together can fail to keep a page out of the index if it is linked from elsewhere. To reliably keep a page out of results, allow crawling and use noindex; use robots.txt to manage crawl efficiency, not to hide pages. Getting this wrong is a common cause of pages that stubbornly remain indexed.

How to add noindex and what to use it on

Noindex is applied with a robots meta tag in a page's HTML head, or an X-Robots-Tag in the HTTP header for non-HTML files. It suits pages you want accessible to users but not in search results: thank-you and confirmation pages, internal search results, thin tag or filter pages, staging or duplicate content, and admin areas. It should not be applied to pages you actually want to rank, and it must not be combined with a robots.txt block on the same page, which would stop Google seeing the tag. Because a wrongly placed noindex silently removes a page from search, it is worth auditing after template changes to ensure no important page carries one by mistake.

FAQ

Will noindex stop Google from crawling a page?

No. Noindex only prevents Google from including a page in its search index. Google may still crawl the page to read the noindex tag. If you want to prevent crawling entirely, you use robots.txt to block the bot. Noindex and robots.txt serve different purposes and should not be confused.

Which pages should South African businesses noindex?

Common candidates for noindex include thank-you pages after form submissions, internal search results pages, login and account pages, staging or draft content, and paginated pages beyond page 1 on thin category pages. Noindexing low-value pages helps consolidate crawl budget towards your important content.

Does noindex remove a page from Google immediately?

No. Google must recrawl the page to see the noindex tag before it drops from results, which can take days to weeks depending on how often the page is crawled. The page must be crawlable, not blocked in robots.txt, for Google to see the tag at all.

Can noindex hurt SEO?

Applied deliberately to low-value pages, noindex helps by keeping thin or duplicate pages out of the index. The danger is applying it by mistake to pages you want to rank, which silently removes them from search, so noindex tags are worth auditing after any template change.

Want a team that knows these metrics cold?

Founder-led digital marketing for South African businesses since 2015. 4.9-star rated, 64+ clients, no long-term contracts.