Noindex is set with a meta robots tag in the page head or an X-Robots-Tag HTTP header. Other meta robots values include nofollow, noarchive and nosnippet.
A crawler must be able to fetch the page to see the directive, so noindexed pages should not also be blocked in robots.txt.
Noindex keeps thin, private or duplicate pages out of search results. Left on by mistake, often after a staging launch, it can remove important pages from search entirely.
VisibilityKit Crawler records the meta robots and X-Robots-Tag values for every URL and flags indexable pages that carry noindex, so a leftover staging directive does not go unnoticed.
Disallow stops crawling. Noindex stops indexing. A page blocked by robots.txt can still appear in results if other sites link to it.
No. The sitemap should list only pages you want indexed.
Whether a page is able to be included in a search engine's index.
A text file at the root of a site that tells crawlers which paths they may and may not fetch.
An HTML link element that tells search engines which URL is the preferred version of a page.
A file listing the URLs you want search engines to discover and index.
A desktop SEO crawler for macOS, Windows and Linux. Find what is broken, assign the work, then verify every fix. Crawl data stays on your machine.
Get Crawler free