Should you block AI crawlers and SEO bots?

4 min read · updated 9 October 2026

Open the access.log of any site and most lines will come not from visitors or attackers but from bots: SEO services building link databases, and crawlers collecting text to train neural networks. On the servers we checked, they made about 23.5 thousand requests from a thousand addresses in one day. This isn't an attack, and whether to block them is the site owner's decision. Here's who you can block without losing anything, and who you shouldn't touch.

How many there are

Numbers from servers running KIPBan log analysis, for one day as of 9 October 2026:

WhoRequestsAddressesServers
Aggressive SEO bots20,0076384
AI training crawlers3,4584015

For comparison, SSH password guessing on all 13 servers made about 19 thousand attempts in the same day. SEO service bots on four servers made more requests than SSH guessing on all thirteen.

Three different kinds of AI bots

The main mistake is blocking everything with “GPT” or “Claude” in the User-Agent. Large companies run different bots for different jobs:

CompanyModel trainingSearch in the assistantFetch on a user's request
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity—PerplexityBotPerplexity-User
Common CrawlCCBot——
Metameta-externalagent—meta-externalfetcher
  • Training crawlers take text into datasets. Block them and your text won't end up in future model versions. It doesn't affect your site's visibility today.
  • AI search bots build the index an assistant uses to answer questions with links to sources. Block them and the site disappears from those answers, which are a growing source of visits.
  • User fetches are a person who asked the assistant to open your page. Blocking them is like turning a visitor away at the door.

Perplexity has a subtlety. According to Perplexity itself, PerplexityBot builds the index for Perplexity's answers rather than collecting training data. Many lists, including the KIPBan pattern, treat it as a crawler. If visibility in Perplexity matters to you, keep that in mind before turning on crawler blocking.

SEO bots

Bots of link analysis and competitive research services: AhrefsBot, SemrushBot, MJ12bot (Majestic), DotBot (Moz), BLEXBot, MegaIndex. ByteDance's Bytespider is often added to the list too. They crawl the whole site, often, and give the owner nothing unless the owner uses those services. If you do, say you check your own site in Ahrefs, don't block that service's bot.

Search engines are a different matter. Googlebot, Bingbot and YandexBot must never be blocked under any settings, or the site drops out of search.

Option 1. robots.txt

Start with robots.txt. All the large companies respect it, and disallowing a company's training crawler doesn't affect the same company's search bot:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: AhrefsBot
Disallow: /

User-agent: SemrushBot
Disallow: /

Google-Extended isn't a separate bot but a robots.txt token. The site is still crawled by regular Googlebot, and the token forbids using the text to train Gemini. You won't see it in the log and can't block it from the log.

The downside of robots.txt is that it's a request. Bots that ignore it keep coming, and polite ones take hours or days to re-read the file.

Option 2. Refusing by User-Agent in nginx

map $http_user_agent $blocked_bot {
    default 0;
    ~*(GPTBot|ClaudeBot|CCBot|meta-externalagent|AhrefsBot|SemrushBot|MJ12bot|DotBot|BLEXBot|MegaIndex|Bytespider) 1;
}

server {
    # ...
    if ($blocked_bot) { return 403; }
}

This works immediately and precisely, but every request still reaches nginx, and bots that get a 403 usually keep knocking.

Option 3. Banning from the log

The toughest option: a fail2ban jail that bans the address at the firewall after the very first request with such a User-Agent. Requests stop reaching the server entirely. Be careful with the list here: one extra term in the regex, say bot without qualification, and Googlebot goes to the ban list.

How KIPBan does it

The KIPBan pattern catalogue has two patterns the owner turns on by choice: “Aggressive SEO bots” and “AI training crawlers”. We don't call them attack protection; they're a setting. Both ban on the first request, and neither touches:

  • search engines: Google, Bing, Yandex;
  • link previews in messengers and social networks;
  • AI search and fetches on a user's request: ChatGPT-User, OAI-SearchBot, Claude-SearchBot, Perplexity-User.

No KIPBan pattern blocks curl, Wget or python-requests: monitoring, webhooks and payment notifications use them.

Log analysis first shows how many of these bots reach the server, and the jail starts in watch mode: it counts whom it would ban and blocks nothing. You see exactly what you'd cut off before you cut it off. One thing to keep in mind: in ban mode the addresses go to the shared KIPBan database, including those caught by these two patterns.

What to choose

  • You don't want your text used for training but want to stay in assistants' answers: robots.txt for the training crawlers, leave AI search bots alone.
  • Bots load the server: ban from the log the SEO bots you don't use and the training crawlers.
  • The site lives on search traffic: nothing that touches Google, Bing or Yandex, and no broad masks like bot or crawler.

Questions

Will my site disappear from ChatGPT if I block GPTBot?

No. GPTBot collects training data. Search in ChatGPT is handled by OAI-SearchBot, and opening pages on a user's request by ChatGPT-User. As long as those aren't blocked, the site can appear in ChatGPT's answers.

Can I block Google-Extended in nginx?

No. Google-Extended isn't a separate bot but a robots.txt token. Pages are crawled by regular Googlebot, and the robots.txt rule only stops them from being used to train Google's models.

Do AI crawlers respect robots.txt?

The bots of OpenAI, Anthropic, Google, Common Crawl and Meta say they do. Some other bots ignore it, and those need a User-Agent refusal or a ban from the log.