robots.txt lists user agents and the paths they are allowed or disallowed from crawling. It can also point to your XML sitemap.
Well-behaved crawlers, including search engines and most AI crawlers, read it before crawling. It is a request, not access control.
A single wrong line can block search engines or AI crawlers from your whole site. It is also where you decide which AI crawlers may read your content.
VisibilityKit Crawler reports URLs blocked by robots.txt during a crawl, and its AI readiness check reports which AI crawlers your robots.txt allows.
No. It stops crawling. Use noindex to keep a page out of the index.
At the root of the host, for example example.com/robots.txt. Each subdomain needs its own.
Whether AI crawlers such as GPTBot, ClaudeBot and PerplexityBot are allowed to fetch your site.
A directive telling search engines not to include a page in their index.
A file listing the URLs you want search engines to discover and index.
The process by which AI platforms discover and index web content for use in generating responses.
Whether a page is able to be included in a search engine's index.
A desktop SEO crawler for macOS, Windows and Linux. Find what is broken, assign the work, then verify every fix. Crawl data stays on your machine.
Get Crawler free