Question 3 of 6 / Allow

Which AI crawlers may read your site.

Training, search and fetches a person asks for are three different things, and the big AI companies give each its own crawler. Here is which is which, and where HubSpot lets you decide.

Sources read on 24 September 2026, next review by 23 December 2026. Nothing here promises rankings or AI citations.

THE SHORT ANSWER

How do you let AI crawlers in, or keep them out, on HubSpot?

Decide by purpose: training, search, and fetches a person asks for. The major vendors name a separate crawler for each, and HubSpot's robots.txt setting lets you allow or block each one; some fetchers that act for a person say robots.txt may not apply.

Robots.txt controls crawling, not indexing. It is a request that the named crawlers say they honour, not a lock.

Robots.txt on HubSpot

01 / On HubSpot

Five steps,
one decision each.

HubSpot's setting was checked in a HubSpot account on 24 September 2026.

  1. 01

    Read the robots.txt you have

    Open yourdomain.com/robots.txt. Note which crawlers it names today and what it allows them.

    Leave with Your starting point

  2. 02

    Open HubSpot's robots.txt setting

    Settings, Content, Pages. Choose a domain, or Default settings for all domains, then the SEO & Crawlers tab and Robots.txt. Save when you are done.

    Leave with The file you can edit

  3. 03

    Write one group per decision

    Allow the search crawlers you want to appear in: OAI-SearchBot, Claude-SearchBot, PerplexityBot. Decide separately on training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot.

    Leave with Search and training decided apart

  4. 04

    Check Search Console's AI features control

    In Search Console, the Search generative AI features control decides whether your site is eligible for Google's AI features. Include is the default.

    Leave with Eligible for AI Overviews and AI Mode

  5. 05

    Keep noindex and robots.txt apart

    A page blocked in robots.txt cannot be crawled, so its noindex is never read. HubSpot's knowledge base says the two methods should not be combined.

    Leave with Each page out the right way

02 / The evidence

What the AI companies
say about their crawlers.

Each card quotes the company's own crawler page.

01 / OPENAI

Search and training are separate at OpenAI.

OpenAI: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." For ChatGPT-User: "Because these actions are initiated by a user, robots.txt rules may not apply."

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why this matters

OpenAI also lists GPTBot for training and OAI-AdsBot, which only visits pages submitted as ads. The page shows no date.

02 / ANTHROPIC

Anthropic says all three of its bots obey robots.txt.

Anthropic: "Anthropic's Bots respect “do not crawl” signals by honoring industry standard directives in robots.txt." That covers ClaudeBot, Claude-SearchBot and Claude-User.

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why this matters

Updated on 7 April 2026. Anthropic makes no exception for the fetcher that acts for a person, unlike OpenAI, Perplexity and Google.

03 / GOOGLE

Google-Extended is not about Search.

Google: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why this matters

For AI Overviews and AI Mode, Google names Googlebot and the snippet controls as the controls. Last updated 14 July 2026.

04 / PERPLEXITY

Perplexity's user fetches ignore robots.txt.

Perplexity, on Perplexity-User: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." PerplexityBot, its search crawler, is "not used to crawl content for AI foundation models."

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why this matters

The page's own data says it was modified on 29 January 2026.

05 / APPLE

Applebot-Extended only opts you out of training.

Apple: "Applebot-Extended does not crawl webpages. Webpages that disallow Applebot-Extended can still be included in search results."

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why this matters

Published 4 September 2026. Without Applebot rules, Apple says its robot follows Googlebot's.

03 / Questions

Fair questions,
straight answers.

What people ask before they touch robots.txt.

If I block Google-Extended, do I leave AI Overviews?

Not according to Google's own pages read together: Google-Extended governs Gemini training and grounding in some of Google's other systems, while for AI Overviews and AI Mode Google names Googlebot and the snippet controls. That is our reading of two pages; Google does not say it in one sentence.

Google's common crawlers

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Can I stop ChatGPT from reading a page someone pastes in?

Not reliably with robots.txt. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says its user fetcher generally ignores it. Google says its user-triggered fetchers "generally ignore robots.txt rules". Anthropic says Claude-User follows robots.txt.

Google's user-triggered fetchers

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Will blocking training crawlers hide me from AI search?

The companies keep the two apart. OpenAI: GPTBot for training, OAI-SearchBot for ChatGPT search. Anthropic: ClaudeBot and Claude-SearchBot. Apple: Applebot-Extended only opts out of training. You can refuse training and still allow search.

OpenAI's crawlers

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Is Perplexity crawling in secret?

Cloudflare said so on 4 August 2025, and Perplexity disputed it the same day. It is a dispute, not a finding. What Perplexity documents is PerplexityBot for search and Perplexity-User for fetches a person asks for.

Cloudflare's post

Owned by Freelopers OÜ
Last reviewed 24 September 2026

Why is my hs-sites.com address not indexed?

HubSpot: "Content on HubSpot system domains containing hs-sites is always set as no-index in a robots.txt file." Connect your own domain and publish there.

HubSpot on robots.txt

Owned by Freelopers OÜ
Last reviewed 24 September 2026

04 / Sources

Sources,
read on 24 September 2026.

Every page we quote, with its publisher and the date it shows. We re-read them at every review.

  1. 01

    Overview of OpenAI crawlers

    OpenAINo date shownDocumentation

    GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot.

  2. 02

    Does Anthropic crawl data from the web?

    Anthropic7 April 2026Documentation

    ClaudeBot, Claude-SearchBot, Claude-User.

  3. 03

    Perplexity crawlers

    Perplexity29 January 2026Documentation

    PerplexityBot and Perplexity-User.

  4. 04

    Google's common crawlers

    Google14 July 2026Documentation

    Google-Extended and Search.

  5. 05

    Google's user-triggered fetchers

    Google19 August 2026Documentation

    Fetchers that act for a person.

  6. 06

    About Applebot

    Apple4 September 2026Documentation

    Applebot and Applebot-Extended.

  7. 07

    Search generative AI features control

    Google Search Console HelpNo date shownDocumentation

    Eligibility for Google's AI features.

  8. 08

    Prevent content from appearing in search results

    HubSpot Knowledge Base28 July 2026Documentation

    Editing robots.txt on HubSpot.

  9. 09

    Perplexity is using stealth, undeclared crawlers

    Cloudflare4 August 2025Article

    Disputed by Perplexity the same day.

05 / Where you are

Next: does llms.txt
help at all?

Or pick the route closest to where you are.

Find out what AI reads today.

Guard checks crawler access, structured data, speed and accessibility every week and tracks whether real AI answers mention you.

Does your brand appear when someone asks about your field?
Run a free Guard audit

Start free with DeLight.

A free HubSpot theme with the site graph and structured data built in, installed more than 1,900 times.

Do you need a site before you need answers?
Get DeLight free

Let the theme do the structure.

Answerable puts the answer first, attaches the evidence and publishes structured data that matches the page, checked in public.

Which of your pages answer a real question?
See the proof page

Have Freelopers do it with you.

We build and fix HubSpot sites for answer engines: the pages, the data, the crawler rules and the measuring.

Would you rather spend the time on your content?
Talk to us