AI visibility · from each company’s own documentation

Can AI see your website?

Eight AI systems might mention a small business. Each company publishes its own rules for its robots — and some of those robots decide whether you can be seen at all, while others only decide whether your pages help train a model. Those are different choices. The “block the AI” lines people copy off the internet mix them up, and a site can end up blocked from a system without anyone having meant it.

Eight companies, eight rulebooksEvery rule quoted from the company itselfFree check, no account
Free · no account, no email required

Can the eight AI systems actually see your website?

Each of these companies publishes rules about which of its robots may read your site. Most small-business websites are blocked from at least one — usually by a line somebody pasted in years ago trying to keep AI out, which blocked the wrong thing.1

Showing a worked example on a demonstration site until you enter your own. Leave an email and we’ll send you the result; there’s no account to make and nothing to cancel. How we handle it.

Google AI Overviews & AI Mode
Search Console control set to Inherit — the parent property says Exclude
blocked
ChatGPT search
OAI-SearchBot disallowed in robots.txt
blocked
Perplexity
PerplexityBot disallowed in robots.txt
blocked
Microsoft Copilot (Bing)
noarchive on every page — does nothing on Google, blocks Copilot on Bing
degraded
Google Gemini apps
Google-Extended disallowed. This one may well be deliberate
your choice
Claude
No rule names any of Anthropic's three robots
clear
Meta AI
No rule names Meta-WebIndexer
clear
Apple — Siri, Spotlight, Safari
No rule names Applebot, and no Googlebot restriction to inherit
clear
Three blocked, one degraded, three clear — and the three that are clear are clear by luck. Nothing in this site’s file names Anthropic’s, Meta’s or Apple’s robots, so those three can see it. Not because anyone decided they should: because the list somebody pasted in didn’t happen to include them.

Two different switches

Most of these companies run more than one robot. One kind reads your site so the company’s AI can find, show and link to you when someone asks a question. Another kind collects pages to train the company’s models. You might reasonably want to say no to training. Almost no business wants to say no to being found.

The trouble is that each company draws that line in a different place, and uses different names. A snippet that blocks “the AI bots” will be right for some companies and wrong for others — it cannot be right for all eight, because the eight don’t agree with each other.

The eight, one at a time

as each company documents it

Google AI Overviews and AI Mode

Decides whether you can be seen: A setting inside Search Console called Search generative AI — not your robots.txt file at all.
Governs training: Google-Extended, which governs Google’s Gemini models, not Search.
Of the Exclude setting: “You won’t receive any traffic or impressions from these features.”2

Google Gemini apps

Decides whether you can be seen: Google documents no separate switch.
Governs training: Google-Extended
“Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”1

Microsoft Copilot

Decides whether you can be seen: Bing’s robot, Bingbot, being allowed in — and no NOARCHIVE or NOCACHE tag on the page.
Governs training: Microsoft documents no separate one.
NOARCHIVE “prevents content from being used in Copilot responses and grounding results.”4

ChatGPT search

Decides whether you can be seen: OAI-SearchBot
Governs training: GPTBot
“Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.”1

Perplexity

Decides whether you can be seen: PerplexityBot
Governs training: None. Perplexity says PerplexityBot “is not used to crawl content for AI foundation models.”
“To ensure your site appears in search results, we recommend allowing PerplexityBot in your site’s robots.txt file.”1

Claude

Decides whether you can be seen: Claude-SearchBot and Claude-User
Governs training: ClaudeBot
“Disabling Claude-SearchBot on your site prevents our system from indexing your content for search optimization, which may reduce your site’s visibility and accuracy in user search results.”1

Meta AI

Decides whether you can be seen: Meta-WebIndexer
Governs training: Meta-ExternalAgent — which Meta says is also used for “improving products by indexing content directly.” Blocking it is not purely a training decision.
“Allowing Meta-WebIndexer in your robots.txt file helps us cite and link to your content in Meta AI’s responses.”1

Apple — Siri, Spotlight and Safari

Decides whether you can be seen: Applebot. Apple also documents a nosnippet setting for keeping a page out of the AI-generated answers it gives to general-knowledge questions in Siri and Search.
Governs training: Applebot-Extended — “Webpages that disallow Applebot-Extended can still be included in search results.”
“Enabling Applebot in robots.txt allows website content to appear in search results for Apple users around the world in these products.”1

Robot names are written exactly as each company publishes them, because that is how your robots.txt file has to spell them. Quotations are the companies’ own words; we read each page and record the date we read it.

Four places the copied snippets go wrong

The robots that fetch a page when someone asks a question are not treated alike.
OpenAI says robots.txt rules “may not apply” to its ChatGPT-User fetcher, and Perplexity says Perplexity-User “generally ignores robots.txt rules.” Anthropic says the opposite about its own: blocking Claude-User “may reduce your site’s visibility for user-directed web search.” On Claude, that robot is a visibility switch.1
A “training” robot is not always only a training robot.
Perplexity separates the two cleanly. Meta does not: Meta-ExternalAgent crawls “for use cases such as training foundation AI models or improving products by indexing content directly.” Blocking it on Meta is not a pure training decision, and shouldn’t be described to you as one.1
One company quietly follows another company’s rules.
Apple: “If robots instructions don’t mention Applebot but mention Googlebot, the Apple robot will follow Googlebot instructions.” Any restriction you placed on Google’s robot is also being applied to Apple’s, whether you knew it or not.1
Even the small print differs.
Anthropic says it supports the Crawl-delay line in robots.txt. Apple says “Applebot does not follow crawl-delay.” A file tuned for one is not tuned for the other.1

The Google switch that isn’t in your robots.txt

Google’s AI Overviews and AI Mode are governed by a setting inside Search Console, under Settings → Search generative AI. Google’s own guide says a site “must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search.”3 As of August 31, 2026, Google says the control has been rolled out to all websites worldwide. It has three settings: include, exclude, or inherit whatever a higher-level property says.2

Include is the default, so a property nobody has touched is fine. The cases worth checking are the ones a small business inherits without knowing: a previous web person who switched it to Exclude, or a higher-level property set to Exclude that your own property quietly follows. Neither can be seen from your website. They can only be read inside Search Console — which is why our free check can’t see them and the full audit reads both.

Two things Google is careful to say, and so are we. The setting “isn’t used as a ranking or inclusion signal affecting other parts of Search.” And switching it back to Include removes a block; it doesn’t put you in the AI answers. It makes you eligible.2

What we won’t sell you

A lot of “AI optimization” for sale right now is things Google’s own guide says it does not use. When someone offers you one, this is what Google wrote:3

An “llms.txt” file
“Doing so will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.”
Rewriting your pages into small “chunks” for AI
“There’s no requirement to break your content into tiny pieces for AI to better understand it.”
Writing in a special way “for the AI”
“You don’t need to write in a specific way just for generative AI search.”
Special AI markup
“Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.”
Hidden instructions for AI robots tucked into your pages
Never, for anyone. Bing’s webmaster guidelines list it as prohibited.4 We won’t do it and we’ll tell you if we find it.

Google’s note covers Google Search only. If another of the eight companies documents a use for the file, we’ll say so — and quote them.

What this can’t tell you

Being seen is not the same as being chosen.

  • Whether any AI will mention you. Every one of the eight publishes what its robots can reach. None publishes how its AI answers choose which sources to cite. Apple comes closest, with an unweighted list of what its search ranking considers.
  • When anything will change. Google documents how long it takes to exclude a site after the setting changes. It gives no timeline for being included again, and neither do we.
  • Anything the documentation doesn’t say. Where a company is silent, we say it’s silent and name where we looked.

What the check does do is find blocks you didn’t mean to put there. Removing those is the part that’s knowable, so it’s the part we sell. “Indexing and serving aren’t guaranteed,” in Google’s words — for AI answers as much as for search results.3

References

  1. Each company’s own crawler documentation. Google, “List of Google’s common crawlers” (updated 2026-07-14); OpenAI, “Overview of OpenAI Crawlers”; Perplexity, “Perplexity Crawlers” — read 2026-09-05 and checked in a browser 2026-09-08. Anthropic, “Does Anthropic crawl data from the web…”; Apple, “About Applebot” (published 2026-09-04); Meta, “Meta Web Crawlers” (updated 2026-05-21) — read 2026-09-15.
  2. Google, “Search generative AI control,” Search Console Help — the three settings, inheritance from a parent property, what exclusion does and does not do, and the time it takes. Read 2026-09-05; checked in a browser with every section expanded 2026-09-08.
  3. Google, “Optimizing your website for generative AI features on Google Search” (updated 2026-07-10) — the eligibility sentence, the llms.txt, chunking, writing and markup sentences, and “Indexing and serving aren’t guaranteed.” Read 2026-09-05.
  4. Bing, Webmaster GuidelinesNOARCHIVE and Copilot, and hidden instructions for AI listed as prohibited. Read from our archived copy of 2026-08-30.

Every company named here publishes its own rules, and you are entitled to check us against them. If anyone tells you how ChatGPT or Google’s AI decides what to show, ask for the page it came from.