AI companies run their own crawlers to gather content for training and for live answers. Each identifies itself with a user agent and generally respects robots.txt.
Access is decided per crawler, so a site can allow search-focused AI crawlers while blocking training crawlers, or the other way round.
If your robots.txt blocks an AI crawler, that platform may not be able to read your pages when it answers questions. Many sites block them by accident through broad rules.
VisibilityKit Crawler's AI readiness check reports which AI crawlers your robots.txt allows and validates your llms.txt, alongside the usual technical audit.
If you want AI assistants to read and cite your content, allow the crawlers they use for search and answers. Blocking training crawlers is a separate choice.
No. It removes a blocker. Whether you are cited has to be measured.
A text file at the root of a site that tells crawlers which paths they may and may not fetch.
The process by which AI platforms discover and index web content for use in generating responses.
A proposed standard file that tells AI crawlers what content is available and how to access it on your website.
The process by which AI platforms catalog and store web content for retrieval during response generation.
Running a page's JavaScript to see the content and links that only appear after scripts execute.
A desktop SEO crawler for macOS, Windows and Linux. Find what is broken, assign the work, then verify every fix. Crawl data stays on your machine.
Get Crawler freeSee how often ChatGPT, Perplexity, Gemini and Google AI Overviews mention, recommend and cite your brand, with the evidence behind every answer.
Explore Intelligence