When an assistant looks something up mid-answer, it sends out its crawler. The crawler reads robots.txt first and checks whether it is allowed in. If it finds a block it leaves, and your site stops existing for that assistant, even though Google works fine.
In the audits we run, blocked AI crawlers are one of the three most common reasons a business is absent from answers. It is almost never a deliberate choice. Usually a plugin, the agency that built the site, or the hosting provider added it.
Who is knocking
| Crawler | Whose | What it does |
|---|---|---|
GPTBot | OpenAI | Fetches pages for model training |
OAI-SearchBot | OpenAI | Search during a ChatGPT answer |
ChatGPT-User | OpenAI | Fetches a page at a user’s explicit request |
PerplexityBot | Perplexity | Indexes pages for cited answers |
ClaudeBot | Anthropic | Fetches pages for Claude |
Google-Extended | Permission for Gemini to use your content | |
Applebot-Extended | Apple | Apple’s equivalent |
Bingbot | Microsoft | Bing search, and the source behind Copilot |
GPTBot collects training data, while OAI-SearchBot handles search during a conversation. If you do not want your content used for training but do want to be recommended, block the first and allow the second.A ready file for a business that wants to be recommended
Paste into robots.txt in your site root. Change the sitemap address to your own.
User-agent: *
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml
If you also want to keep your content out of training data, add at the end:
User-agent: GPTBot
Disallow: /
How to check what you have now
- Type
yourdomain.com/robots.txtinto your browser. The file shows as plain text. - Look for the line
Disallow: /. If it sits underUser-agent: *, you are blocking everything and everyone, Google included. - Check whether any crawler from the table above has a
Disallownext to it. - Look in your page source for
<meta name="robots">. A value ofnoindexworks independently of robots.txt. - If you use Cloudflare, check the AI crawler blocking setting in the dashboard. It is sometimes on by default and operates above robots.txt.
That last point is the sneaky one. Cloudflare can block AI crawlers at the network level, so robots.txt may look correct while the crawler still never gets in. If you fixed the file and nothing changed, look there.
The most common accidental blocks
- A staging site that went live. The developer blocked everything during the build and forgot to lift it at launch.
- A privacy plugin. Some add AI crawler blocks as "content protection", without asking.
- A WordPress setting. The "Discourage search engines from indexing this site" box under Reading. Once ticked, it can survive for years.
- A file copied from another site. Along with blocks that had nothing to do with yours.
What comes after the crawlers get in
Letting crawlers in is necessary, not sufficient. The crawler arrives and has to read something. If your site is one image of a menu and a contact form, getting in changes nothing.
The next two steps are LocalBusiness structured data, so nobody has to guess the basic facts, and a tidy Google Business Profile, where the model gets your category, hours and reviews. It all adds up to the approach described on the Method page.
And if you want to see where you stand first, order the free scan. We check whether your robots.txt blocks assistants along the way, and say so in the report.
