Technical

Should I block AI crawlers on my website?

The answer depends on what you sell. For a publisher, blocking can make sense. For a bakery it is an own goal. The good news is that the two decisions can be separated.

·5 min read
Short answer. If content is your product, block the training crawler. If your product is a service you deliver on site, block nothing, because you are cutting yourself out of recommendations. You can also do both at once.

Two different things that are easy to confuse

The debate about AI crawlers mixes up two matters with very different consequences.

What the crawler doesWhoWhat it means for you
Fetches pages to train a modelGPTBot, ClaudeBot, Google-ExtendedYour content enters the model permanently, with no link and no traffic back
Searches while answeringOAI-SearchBot, PerplexityBot, ChatGPT-UserYour business can be named in an answer to a customer’s question

Blocking the first group protects content. Blocking the second removes you from recommendations. They are not the same decision and do not have to be made together.

When blocking makes sense

  • Publisher or author. You sell articles, courses, reports. A model that absorbs the lot starts answering instead of you.
  • A database as the product. A directory, a search tool, a price comparison. Your value lies in the completeness of the set.
  • Content under contract. Material you do not fully own or that is licensed against further use.

When blocking hurts

  • A service delivered on site. Restaurant, garage, clinic, salon. The customer is not buying your text, they are turning up at an address.
  • Local retail. A shop you have to be in to buy from.
  • Any business that wants to be recommended. Which is most of them.

For that group blocking is a cost with no benefit. You protect a description of your dining room from a model, and in exchange you lose the chance that it points somebody looking for dinner at your address.

How to do one without the other

Put an allow for the search crawlers and a block for the training one into robots.txt:

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

The full crawler list and a ready file for a business that wants to be recommended are in the post on GPTBot and robots.txt.

The trap: a block you never set

In audits we find accidental blocks more often than deliberate ones. The usual sources are a privacy plugin, a WordPress setting left over from the build, and the AI crawler blocking option in the Cloudflare dashboard, which operates above robots.txt.

If you fixed the file and nothing changed, check Cloudflare. It is the most common reason a correct fix appears not to work.

Not sure where you stand? The free scan checks your robots.txt along the way and says in the report whether anything is blocking assistants.

Find out whether the assistants name your business

We ask ChatGPT three of your customers’ questions and show you the full answers with timestamps. You get the result straight away, no card.

Get the free scan
Free scan