Is your website blocking AI assistants? robots.txt and Cloudflare, explained for practices
Many practice websites turn away ChatGPT, Claude, and Perplexity without anyone deciding to. How to check yours in ten minutes, and what to change.
Before an AI assistant can recommend your practice, it has to be able to read your website. A surprising number of practice sites quietly turn those readers away. Nobody decided to; a template, plugin, or security setting did it for them.
There are two places this happens: a small file on your site called robots.txt, and the security service in front of your site, if you have one.
What robots.txt does
Every website can have a plain-text file at /robots.txt (for example, yourpractice.com/robots.txt). It lists automated readers, called crawlers or bots, by name and says which parts of the site each may visit. Well-behaved crawlers, including the ones run by the major AI companies, follow it.
A rule like this tells one crawler to stay out of the whole site:
User-agent: GPTBot
Disallow: /
And this tells every crawler to stay out, which also blocks Google and Bing:
User-agent: *
Disallow: /
That second version sometimes survives from when a site was being built and was never meant to go live.
The crawlers that matter
Each AI company runs one or more crawlers, and they do different jobs. Some collect pages to train future models. Others fetch pages to answer a question right now, which is the kind that can lead to your practice being named.
- OpenAI:
GPTBot(training),OAI-SearchBot(ChatGPT search),ChatGPT-User(fetching a page for a user) - Anthropic:
ClaudeBot(training),Claude-SearchBot(search),Claude-User(fetching for a user) - Perplexity:
PerplexityBot(search),Perplexity-User(fetching for a user) - Google:
Googlebot(Google Search, including AI Overviews),Google-Extended(controls use in Gemini; it doesn't affect Google Search) - Microsoft:
Bingbot(Bing and Copilot; other assistants also draw on Bing results) - Apple:
Applebot-Extended(Apple Intelligence)
You can allow search and leave training blocked if you prefer. That's a reasonable choice. What hurts is blocking the search crawlers without meaning to.
Check yours in ten minutes
- Open your robots.txt. Type your web address followed by
/robots.txtinto a browser. If you get a "page not found", that's fine: no file means nothing is blocked. - Look for
Disallow: /. Note whichUser-agentline it sits under. Under*, it blocks everyone. Under a crawler from the list above, it blocks that one. - Check your website builder's settings. Some builders, Squarespace for example, have a setting that blocks known AI crawlers. Look in the SEO or crawler settings.
- Ask whether you use Cloudflare. If your web person or IT company set up Cloudflare (often for speed or security), read the next section.
- Test it. Ask ChatGPT or Perplexity to summarize a specific page on your site by pasting its address. If it says it can't access the page, something is blocking it.
The Cloudflare catch
Cloudflare sits in front of a large share of websites and filters traffic before it reaches them. Cloudflare now blocks many AI crawlers by default on newly added sites, and its bot-protection settings can block others.
This happens before robots.txt is ever read. So your robots.txt can say every crawler is welcome while Cloudflare is still turning them away. If your site uses Cloudflare, have whoever manages it open the domain's bot settings and confirm AI crawlers aren't blocked, or are allowed for the ones you want.
What to change
If you find a block you didn't intend, the fix is usually small:
- robots.txt: remove the
Disallow: /line for the crawlers you want, or add an explicitAllow: /for them. Keep any rules that protect private areas like patient portals. - Website builder: turn off the AI-crawler block, or adjust it to allow search crawlers.
- Cloudflare: allow the AI crawlers you want in the bot settings.
Then check again in a week. Crawlers revisit on their own schedule, so changes aren't instant.
AdLivo's AI Visibility page runs these checks for you. Its deeper website scan reads your robots.txt, notes which AI crawlers are blocked, flags bot checks like Cloudflare's, and gives you a ready-to-use robots.txt you can hand to your web person.
