PerplexityBot

Disputed robots.txt compliance

PerplexityBot crawls and indexes content specifically to power Perplexity's cited, synthesized answers — a live-retrieval crawler rather than a training one, which makes crawl access unusually important for this engine specifically.

User-agent
PerplexityBot
Operator
Perplexity AI

Allow in robots.txt

User-agent: PerplexityBot
Allow: /

Block in robots.txt

User-agent: PerplexityBot
Disallow: /

Perplexity states PerplexityBot honors robots.txt. Independent reporting in 2024 raised questions about fetching through undeclared infrastructure, so some publishers additionally block by IP range as a precaution.

Why crawl access matters more for Perplexity than most engines

Perplexity leans harder on real-time web retrieval than most AI answer engines, which means PerplexityBot access is a structural prerequisite for citation on that engine specifically, not just a helpful signal. If PerplexityBot is blocked, you're not underweighted for a given query — you're excluded from being cited on it entirely, no matter how strong the underlying content is. This makes checking PerplexityBot access a higher-priority item than it might be for engines that lean more heavily on training data.

Frequently asked questions

Yes — reporting has questioned whether some content was fetched outside the declared PerplexityBot user-agent. Perplexity disputes that this bypasses robots.txt for the named crawler.

Verify your robots.txt against every AI crawler

Scoutern's free AI bot checker tests exactly which bots can reach your site.