CCBot
Honors robots.txtCCBot crawls the web to build the open, freely-downloadable Common Crawl dataset, which many AI labs use as raw pretraining data, either directly or via datasets derived from it.
- User-agent
- CCBot
- Operator
- Common Crawl (nonprofit)
Allow in robots.txt
User-agent: CCBot
Allow: /Block in robots.txt
User-agent: CCBot
Disallow: /CCBot isn't tied to one named AI vendor — blocking it reduces your presence in a dataset used by many different labs and open-source model builders, not just one company.