Technical

GPTBot, PerplexityBot and the rest: what to put in robots.txt

One line in a text file can shut your site to every assistant at once. Checking takes thirty seconds, and without it none of the other work matters.

·6 min read

When an assistant looks something up mid-answer, it sends out its crawler. The crawler reads robots.txt first and checks whether it is allowed in. If it finds a block it leaves, and your site stops existing for that assistant, even though Google works fine.

In the audits we run, blocked AI crawlers are one of the three most common reasons a business is absent from answers. It is almost never a deliberate choice. Usually a plugin, the agency that built the site, or the hosting provider added it.

Who is knocking

CrawlerWhoseWhat it does
GPTBotOpenAIFetches pages for model training
OAI-SearchBotOpenAISearch during a ChatGPT answer
ChatGPT-UserOpenAIFetches a page at a user’s explicit request
PerplexityBotPerplexityIndexes pages for cited answers
ClaudeBotAnthropicFetches pages for Claude
Google-ExtendedGooglePermission for Gemini to use your content
Applebot-ExtendedAppleApple’s equivalent
BingbotMicrosoftBing search, and the source behind Copilot
Note the split at OpenAI. GPTBot collects training data, while OAI-SearchBot handles search during a conversation. If you do not want your content used for training but do want to be recommended, block the first and allow the second.

Paste into robots.txt in your site root. Change the sitemap address to your own.

User-agent: *
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

If you also want to keep your content out of training data, add at the end:

User-agent: GPTBot
Disallow: /

How to check what you have now

  1. Type yourdomain.com/robots.txt into your browser. The file shows as plain text.
  2. Look for the line Disallow: /. If it sits under User-agent: *, you are blocking everything and everyone, Google included.
  3. Check whether any crawler from the table above has a Disallow next to it.
  4. Look in your page source for <meta name="robots">. A value of noindex works independently of robots.txt.
  5. If you use Cloudflare, check the AI crawler blocking setting in the dashboard. It is sometimes on by default and operates above robots.txt.

That last point is the sneaky one. Cloudflare can block AI crawlers at the network level, so robots.txt may look correct while the crawler still never gets in. If you fixed the file and nothing changed, look there.

The most common accidental blocks

  • A staging site that went live. The developer blocked everything during the build and forgot to lift it at launch.
  • A privacy plugin. Some add AI crawler blocks as "content protection", without asking.
  • A WordPress setting. The "Discourage search engines from indexing this site" box under Reading. Once ticked, it can survive for years.
  • A file copied from another site. Along with blocks that had nothing to do with yours.

What comes after the crawlers get in

Letting crawlers in is necessary, not sufficient. The crawler arrives and has to read something. If your site is one image of a menu and a contact form, getting in changes nothing.

The next two steps are LocalBusiness structured data, so nobody has to guess the basic facts, and a tidy Google Business Profile, where the model gets your category, hours and reviews. It all adds up to the approach described on the Method page.

And if you want to see where you stand first, order the free scan. We check whether your robots.txt blocks assistants along the way, and say so in the report.

Find out whether the assistants name your business

We ask ChatGPT three of your customers’ questions and show you the full answers with timestamps. You get the result straight away, no card.

Get the free scan
Free scan