Blocking GPTBot and ClaudeBot stops model training, not AI citations. Here is what each crawler does and the robots.txt most operators actually want.
Originally published August 23, 2026
The question almost always arrives framed backwards. Operators ask "should I block GPTBot and ClaudeBot to protect my content," when the decision that actually moves their AI visibility is which crawlers they allow. GPTBot and ClaudeBot are training crawlers. Blocking them keeps your pages out of future model training runs. It does not pull you out of ChatGPT or Claude answers today, because live citations run through a different set of bots entirely.
The confusion is understandable, because OpenAI and Anthropic each run three distinct crawlers that most people lump into one name. Get the split right and the tradeoff becomes obvious. Signals runs an aged Reddit account marketplace plus an editorial network for AI brand mentions across Reddit, Quora, Product Hunt, and Threads, and the pattern we see across client audits is consistent: crawl policy is a small lever, and the durable one is being cited in the sources these models retrieve from.
The single most important distinction is training versus retrieval. Training crawlers collect content to improve future model versions. Search and live-fetch crawlers pull pages to answer a query right now, and those are the visits that turn into citations. OpenAI documents GPTBot as the crawler "used to make our generative AI foundation models more useful and safe," while OAI-SearchBot is "used to surface websites in search results in ChatGPT's search features." Anthropic mirrors the structure with ClaudeBot for training, Claude-SearchBot for its search index, and Claude-User for live retrieval when someone asks Claude a question.
| Crawler | Operator | Job | Blocking it costs you | Honors robots.txt |
|---|---|---|---|---|
GPTBot | OpenAI | Model training | Future training inclusion | Yes |
OAI-SearchBot | OpenAI | ChatGPT search index | ChatGPT search citations | Yes |
ChatGPT-User | OpenAI | Live user-triggered fetch | Little reliably (see below) | No |
ClaudeBot | Anthropic | Model training | Future training inclusion | Yes |
Claude-SearchBot | Anthropic | Claude search index | Claude search citations | Yes |
Claude-User | Anthropic | Live user-triggered fetch | Live retrieval into answers | Yes |
Google-Extended | Gemini training and grounding | Gemini eligibility only | Yes | |
Googlebot | Search index (feeds AI Overviews) | Search and AI Overviews | Yes |
Less than most people fear, and less than most people hope. Blocking GPTBot or ClaudeBot removes your pages from the next training corpus, so a future model will know less about your brand from your own site. It changes nothing about whether you get cited this week, because those citations come from the search and live-fetch bots. The catch is that training inclusion is where long-term, unprompted brand knowledge is built, and once a cutoff passes you cannot retroactively add yourself.
There is a second reason blocking training crawlers rarely pays off for a normal brand: leverage. When Reddit signed a reported $60M-per-year licensing deal with Google for training rights, it had leverage most sites will never have. If you are a SaaS company or a DTC brand, blocking GPTBot does not create a negotiation. It just quietly subtracts you from the corpus while your competitors stay in.
This is the expensive mistake. OAI-SearchBot and Claude-SearchBot build the indexes that ChatGPT and Claude pull from when they cite live sources, and Claude-User fetches your page in real time when a user's question points at it. Block those and you are not protecting content, you are deleting yourself from the answer. As Ahrefs has documented, there is only a small overlap between classic Google rankings and AI citations, so search-crawler access is often the only path a model has to find you at answer time.
One caveat that trips people up: ChatGPT-User does not reliably obey robots.txt. OpenAI states that because those actions are user-initiated, "robots.txt rules may not apply." Anthropic takes the stricter line that all its bots honor robots.txt. So a robots.txt block on live-fetch behaves differently across vendors, and it is not a dependable privacy control regardless.
Google-Extended is the one token where blocking is genuinely low-cost, because Google decoupled AI use from Search. It is a robots.txt product token with no separate user agent of its own, and Google states plainly that it "does not affect a site's inclusion or ranking in Google Search." What it governs is eligibility for Gemini uses: future Gemini training, grounding in Gemini apps, and grounding on Vertex AI.
The operator implication is specific. Blocking Google-Extended does not remove you from AI Overviews, because AI Overviews are assembled from the standard Search index that Googlebot feeds. If you want out of Gemini's generative uses but want to keep Search and AI Overviews visibility, Disallow Google-Extended and leave Googlebot alone. That is the cleanest opt-out in the entire crawler landscape, and the only one that carries almost no visibility penalty. If Gemini answers are a channel you care about, weigh the opt-out against how Gemini picks its sources before you disallow the token.
Match the policy to your business model, not to a general unease about scraping. The infrastructure has moved this way already: Cloudflare began blocking AI crawlers by default in July 2025 and launched a Pay Per Crawl marketplace, then in 2026 shifted to Pay Per Use, which pays publishers when their content is actually used in an answer rather than when a bot fetches the page. That model rewards sites with licensing leverage and does little for everyone else.
Block training crawlers if: you are a publisher with paywalled or licensable content, you have negotiating leverage, or you are pursuing per-use compensation. Roughly a quarter of the top 1,000 sites now block GPTBot, up from about 5% in early 2023, and they are overwhelmingly news and media operators.
Allow everything if: you are a brand, SaaS, startup, or creator whose goal is to be found. AI-referred visitors convert at around 4.4x the rate of standard organic traffic per Semrush, even though that traffic is still only about 1% of total visits today. You want to be in the answer.
Block Google-Extended only if: you specifically object to Gemini training and grounding. It has no Search or AI Overviews cost.
So the honest answer to whether you should block GPTBot and ClaudeBot is narrow: do it only when protecting content is worth more to you than being found, and never block the search or live-fetch crawlers by accident while you are at it. Whichever policy you pick, remember that crawl access is table stakes, not strategy. Being allowed to be cited is not the same as being cited. What decides that is how often your brand appears in the third-party sources these engines actually retrieve, which is a separate playbook covered in our guide on how to get mentioned by ChatGPT. For the fuller picture of how much these bots extract versus send back, see our AI crawler traffic breakdown.
No. GPTBot is a training crawler. ChatGPT's live citations come from OAI-SearchBot and ChatGPT-User. Blocking GPTBot only opts you out of future OpenAI model training, not out of ChatGPT answers today.
Blocking Google-Extended will not, because Google states it does not affect inclusion or ranking in Google Search. Blocking Googlebot itself will, and it also removes you from AI Overviews, which are built from the Search index.
No. Robots.txt is a preference that compliant bots honor and others ignore, and OpenAI notes that user-initiated ChatGPT-User fetches may not apply robots.txt rules. Use authentication or noindex for anything that must stay private.
Allow the search and live-fetch crawlers from every engine, and decide on training crawlers based on whether you have a licensing strategy. Then invest in earning brand mentions across the sources these models retrieve from.
:::
Blocking GPTBot and ClaudeBot stops model training, not AI citations. Here is what each crawler does and the robots.txt most operators actually want.
Originally published August 23, 2026
The question almost always arrives framed backwards. Operators ask "should I block GPTBot and ClaudeBot to protect my content," when the decision that actually moves their AI visibility is which crawlers they allow. GPTBot and ClaudeBot are training crawlers. Blocking them keeps your pages out of future model training runs. It does not pull you out of ChatGPT or Claude answers today, because live citations run through a different set of bots entirely.
Blocking GPTBot and ClaudeBot only opts you out of model training. It does not remove you from ChatGPT or Claude answers, which are served by separate search and live-fetch crawlers. Block training bots only if you have a licensing or paywall strategy. If you want AI citations, allow the search and retrieval crawlers and compete on where your brand is mentioned across the web.
The confusion is understandable, because OpenAI and Anthropic each run three distinct crawlers that most people lump into one name. Get the split right and the tradeoff becomes obvious. Signals runs an aged Reddit account marketplace plus an editorial network for AI brand mentions across Reddit, Quora, Product Hunt, and Threads, and the pattern we see across client audits is consistent: crawl policy is a small lever, and the durable one is being cited in the sources these models retrieve from.
The single most important distinction is training versus retrieval. Training crawlers collect content to improve future model versions. Search and live-fetch crawlers pull pages to answer a query right now, and those are the visits that turn into citations. OpenAI documents GPTBot as the crawler "used to make our generative AI foundation models more useful and safe," while OAI-SearchBot is "used to surface websites in search results in ChatGPT's search features." Anthropic mirrors the structure with ClaudeBot for training, Claude-SearchBot for its search index, and Claude-User for live retrieval when someone asks Claude a question.
| Crawler | Operator | Job | Blocking it costs you | Honors robots.txt |
|---|---|---|---|---|
GPTBot | OpenAI | Model training | Future training inclusion | Yes |
OAI-SearchBot | OpenAI | ChatGPT search index | ChatGPT search citations | Yes |
ChatGPT-User | OpenAI | Live user-triggered fetch | Little reliably (see below) | No |
ClaudeBot | Anthropic | Model training | Future training inclusion | Yes |
Claude-SearchBot | Anthropic | Claude search index | Claude search citations | Yes |
Claude-User | Anthropic | Live user-triggered fetch | Live retrieval into answers | Yes |
Google-Extended | Gemini training and grounding | Gemini eligibility only | Yes | |
Googlebot | Search index (feeds AI Overviews) | Search and AI Overviews | Yes |
Less than most people fear, and less than most people hope. Blocking GPTBot or ClaudeBot removes your pages from the next training corpus, so a future model will know less about your brand from your own site. It changes nothing about whether you get cited this week, because those citations come from the search and live-fetch bots. The catch is that training inclusion is where long-term, unprompted brand knowledge is built, and once a cutoff passes you cannot retroactively add yourself.
There is a second reason blocking training crawlers rarely pays off for a normal brand: leverage. When Reddit signed a reported $60M-per-year licensing deal with Google for training rights, it had leverage most sites will never have. If you are a SaaS company or a DTC brand, blocking GPTBot does not create a negotiation. It just quietly subtracts you from the corpus while your competitors stay in.
This is the expensive mistake. OAI-SearchBot and Claude-SearchBot build the indexes that ChatGPT and Claude pull from when they cite live sources, and Claude-User fetches your page in real time when a user's question points at it. Block those and you are not protecting content, you are deleting yourself from the answer. As Ahrefs has documented, there is only a small overlap between classic Google rankings and AI citations, so search-crawler access is often the only path a model has to find you at answer time.
One caveat that trips people up: ChatGPT-User does not reliably obey robots.txt. OpenAI states that because those actions are user-initiated, "robots.txt rules may not apply." Anthropic takes the stricter line that all its bots honor robots.txt. So a robots.txt block on live-fetch behaves differently across vendors, and it is not a dependable privacy control regardless.
Do not use robots.txt as a content firewall. It is a crawl-preference signal that compliant bots respect and non-compliant scrapers ignore. If a page must stay private, put it behind authentication or noindex, not behind a Disallow line.
Google-Extended is the one token where blocking is genuinely low-cost, because Google decoupled AI use from Search. It is a robots.txt product token with no separate user agent of its own, and Google states plainly that it "does not affect a site's inclusion or ranking in Google Search." What it governs is eligibility for Gemini uses: future Gemini training, grounding in Gemini apps, and grounding on Vertex AI.
The operator implication is specific. Blocking Google-Extended does not remove you from AI Overviews, because AI Overviews are assembled from the standard Search index that Googlebot feeds. If you want out of Gemini's generative uses but want to keep Search and AI Overviews visibility, Disallow Google-Extended and leave Googlebot alone. That is the cleanest opt-out in the entire crawler landscape, and the only one that carries almost no visibility penalty. If Gemini answers are a channel you care about, weigh the opt-out against how Gemini picks its sources before you disallow the token.
Match the policy to your business model, not to a general unease about scraping. The infrastructure has moved this way already: Cloudflare began blocking AI crawlers by default in July 2025 and launched a Pay Per Crawl marketplace, then in 2026 shifted to Pay Per Use, which pays publishers when their content is actually used in an answer rather than when a bot fetches the page. That model rewards sites with licensing leverage and does little for everyone else.
Block training crawlers if: you are a publisher with paywalled or licensable content, you have negotiating leverage, or you are pursuing per-use compensation. Roughly a quarter of the top 1,000 sites now block GPTBot, up from about 5% in early 2023, and they are overwhelmingly news and media operators.
Allow everything if: you are a brand, SaaS, startup, or creator whose goal is to be found. AI-referred visitors convert at around 4.4x the rate of standard organic traffic per Semrush, even though that traffic is still only about 1% of total visits today. You want to be in the answer.
Block Google-Extended only if: you specifically object to Gemini training and grounding. It has no Search or AI Overviews cost.
So the honest answer to whether you should block GPTBot and ClaudeBot is narrow: do it only when protecting content is worth more to you than being found, and never block the search or live-fetch crawlers by accident while you are at it. Whichever policy you pick, remember that crawl access is table stakes, not strategy. Being allowed to be cited is not the same as being cited. What decides that is how often your brand appears in the third-party sources these engines actually retrieve, which is a separate playbook covered in our guide on how to get mentioned by ChatGPT. For the fuller picture of how much these bots extract versus send back, see our AI crawler traffic breakdown.
No. GPTBot is a training crawler. ChatGPT's live citations come from OAI-SearchBot and ChatGPT-User. Blocking GPTBot only opts you out of future OpenAI model training, not out of ChatGPT answers today.
Blocking Google-Extended will not, because Google states it does not affect inclusion or ranking in Google Search. Blocking Googlebot itself will, and it also removes you from AI Overviews, which are built from the Search index.
No. Robots.txt is a preference that compliant bots honor and others ignore, and OpenAI notes that user-initiated ChatGPT-User fetches may not apply robots.txt rules. Use authentication or noindex for anything that must stay private.
Allow the search and live-fetch crawlers from every engine, and decide on training crawlers based on whether you have a licensing strategy. Then invest in earning brand mentions across the sources these models retrieve from.
:::
Allowing the crawlers is table stakes. Getting cited is the work. Signals places editorial brand mentions across a 20,000+ site network so the sources AI engines retrieve from actually name you.
Sources