Two different things that are easy to confuse
The debate about AI crawlers mixes up two matters with very different consequences.
| What the crawler does | Who | What it means for you |
|---|---|---|
| Fetches pages to train a model | GPTBot, ClaudeBot, Google-Extended | Your content enters the model permanently, with no link and no traffic back |
| Searches while answering | OAI-SearchBot, PerplexityBot, ChatGPT-User | Your business can be named in an answer to a customer’s question |
Blocking the first group protects content. Blocking the second removes you from recommendations. They are not the same decision and do not have to be made together.
When blocking makes sense
- Publisher or author. You sell articles, courses, reports. A model that absorbs the lot starts answering instead of you.
- A database as the product. A directory, a search tool, a price comparison. Your value lies in the completeness of the set.
- Content under contract. Material you do not fully own or that is licensed against further use.
When blocking hurts
- A service delivered on site. Restaurant, garage, clinic, salon. The customer is not buying your text, they are turning up at an address.
- Local retail. A shop you have to be in to buy from.
- Any business that wants to be recommended. Which is most of them.
For that group blocking is a cost with no benefit. You protect a description of your dining room from a model, and in exchange you lose the chance that it points somebody looking for dinner at your address.
How to do one without the other
Put an allow for the search crawlers and a block for the training one into robots.txt:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
The full crawler list and a ready file for a business that wants to be recommended are in the post on GPTBot and robots.txt.
The trap: a block you never set
In audits we find accidental blocks more often than deliberate ones. The usual sources are a privacy plugin, a WordPress setting left over from the build, and the AI crawler blocking option in the Cloudflare dashboard, which operates above robots.txt.
If you fixed the file and nothing changed, check Cloudflare. It is the most common reason a correct fix appears not to work.
Not sure where you stand? The free scan checks your robots.txt along the way and says in the report whether anything is blocking assistants.
