PricingSearch articles
Book a strategy call
Link Building

9 of 161 SaaS Leaders Block AI Crawlers, None Block AI Search

We read the robots.txt, llms.txt and homepage schema of 164 leading SaaS companies. 9 block an AI crawler, none block AI search, and 63% publish llms.txt.

Editorial graphic of AI search crawlers passing a robots.txt gateway while training crawlers are turned aside
Jordan Ellis October 6, 2026 8 min read 1,508 words

Ask a SaaS marketing team whether they let AI crawlers in. Most can’t answer without opening a file nobody has touched in a year.

So we opened it for them. On 6 October 2026 we read the robots.txt of 164 leading cloud and SaaS companies.

9 of 161 block at least one AI crawler, and none of them block the crawlers behind AI search answers. Every block we found targets model training. 127 of the 161 files don’t mention AI crawlers at all.

The sample is the Forbes Cloud 100 list plus the 65 companies in the BVP Nasdaq Emerging Cloud Index. You can download every row at the end.

The Headline Numbers

Blocking is rare, and llms.txt is common. Here’s the full set.

Signal Companies Share
robots.txt names an AI crawler 34 of 161 21.1%
robots.txt blocks a training crawler 9 of 161 5.6%
robots.txt blocks a search or user-fetch agent 0 of 161 0%
Opts out of AI training by any method 14 of 161 8.7%
Publishes an llms.txt file 102 of 161 63.4%
Organization schema on the homepage 118 of 155 76.1%
A page at /pricing 107 of 158 67.7%
A price in the pricing page HTML 66 of 158 41.8%

The denominators differ because a few sites refused to be read. The method section lists each one.

What SaaS Robots.txt Files Say About AI Crawlers

Most say nothing. 127 of 161 files (78.9%) carry no rule for any of the 16 AI user agents we checked.

Bar chart: 127 of 161 SaaS robots.txt files don't mention AI crawlers, 25 name them and block none, 9 block a training crawler, none block a search agent

Silence means access. A crawler with no rule of its own follows the default group, and no company in the sample closes its whole site by default.

25 companies (15.5%) name AI crawlers only to let them in. The default already allows them, so the line changes nothing technically. It reads as a statement of intent.

Who blocks, and what

9 companies fully disallow at least one AI crawler: Canva, Figma, Forter, Midjourney, Notion, Postman, ServiceNow, Together AI and nCino.

3 block GPTBot (Canva, Midjourney and Figma). 2 block ClaudeBot. 3 block Google-Extended, Google’s opt-out token for Gemini.

The most-blocked crawler is Bytespider, ByteDance’s, with 6 blocks. CCBot, from Common Crawl, has 5.

Postman’s file gives its reason in a comment: “High volume, no citation, no traffic back.”

Training Is Blocked, Search Is Not

AI companies run separate crawlers for separate jobs. Each one can be controlled on its own line.

OpenAI’s crawler documentation lists GPTBot for training, OAI-SearchBot for search results, and ChatGPT-User for pages a person asks it to open. Anthropic’s crawler guidance and Perplexity’s crawler documentation describe the same split.

Bar chart of 16 AI crawlers showing how many SaaS robots.txt files name each one and how many block it

Not one of the 161 files fully blocks a search or user-fetch agent. OAI-SearchBot is named in 23 files, ChatGPT-User in 25 and PerplexityBot in 29. None of them shut it out.

The 9 blockers all drew the line in the same place. They refuse to be training data, and they still want to appear when a buyer asks an assistant for a tool.

For a software company that’s the sensible position. Blocking a search agent removes you from the answer and leaves your competitors in it.

The one company that rations search agents

Figma is the exception worth knowing. It doesn’t block AI search agents, and it doesn’t give them the run of the site.

Its robots.txt disallows everything for five search and user-fetch agents, then lists about 9,000 exact URLs they may read. That’s an allow-list, and it’s the only one in the sample.

Content-Signal lines say the same thing another way

17 files (10.6%) carry a Content-Signal line. It’s a robots.txt extension from Cloudflare’s Content Signals Policy that states a preference for three uses: search, AI input and AI training.

11 of them set ai-train=yes and 5 set ai-train=no. Forter sets both, for different crawler groups.

Add those to the robots.txt blocks and 14 of 161 companies (8.7%) opt out of AI training one way or the other.

llms.txt Adoption Among SaaS Leaders

102 of 161 companies (63.4%) publish an llms.txt file. That’s more than eleven times the number that block anything.

llms.txt is a proposed plain-text file listing a site’s most useful pages for language models, set out in the llms.txt proposal. The median file in our sample is about 13 KB.

Public companies are ahead. 46 of 65 index constituents have one (70.8%), against 57 of 97 Cloud 100 companies (58.8%).

Adoption isn’t proof of effect. Publishing the file shows intent, and it doesn’t show that any assistant reads it. Our guide covers what llms.txt is and what it can’t do.

We also logged who reads llms.txt on our own site.

What a Crawler Can Read on the Page

Access is only half of it. We also checked the raw HTML each server returns, which is what a crawler gets if it doesn’t run JavaScript.

Bar chart of the share of SaaS leaders publishing Organization schema, llms.txt, pricing in HTML and AI crawler rules

118 of 155 readable homepages (76.1%) carry Organization schema. 111 (71.6%) include sameAs links to official profiles. Those two tell a machine who the company is.

Product-level markup is rarer. 27 homepages (17.4%) carry Product or SoftwareApplication schema, and 6 carry FAQPage.

Pricing is the gap

107 of 158 companies (67.7%) have a page at /pricing. Only 66 (41.8%) show a currency amount in that page’s HTML.

The other 41 pricing pages load the number with JavaScript or don’t publish one. A crawler reading the HTML finds no price on either kind.

When the vendor’s own page has no number, an answer about cost has to come from somewhere else.

What to Do With This

Check your own file before anything else. Access isn’t a way to get ahead, because nearly every leading SaaS company already allows AI search crawlers.

  1. Open your robots.txt and search for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot.
  2. Remove any rule that disallows a search or user-fetch agent. No company in this sample has one.
  3. Block the training crawlers by name if you want out of training, and leave the search agents alone.
  4. Load your pricing page with JavaScript off. A number that disappears is a number a crawler can’t read.

One caveat: robots.txt is a request, and a firewall rule can still block a crawler the file allows. 8 sites here refused our own script.

Your server logs show who gets through. We cover how to check which AI bots crawl your site.

Once the door is open, what decides whether you’re named is what other sites say about you. That’s the work behind GEO for SaaS brands.

How We Collected the Data

The sample is 164 companies: the 2025 Forbes Cloud 100 and the 65 constituents of the BVP Nasdaq Emerging Cloud Index on 6 October 2026. Netskope is on both lists and counted once.

For each company we requested four URLs on the host its homepage resolves to: /robots.txt, /llms.txt, the homepage and /pricing.

We parsed robots.txt by the group and longest-match rules in RFC 9309, the robots.txt standard. A crawler counts as blocked when its own group disallows the homepage and four other common paths.

The 16 user agents are PerplexityBot, ClaudeBot, ChatGPT-User, GPTBot, OAI-SearchBot, Google-Extended, CCBot, Applebot-Extended, Perplexity-User, Meta-ExternalAgent, Claude-SearchBot, anthropic-ai, Bytespider, cohere-ai, Claude-User and Amazonbot.

An llms.txt counts when the server returns a text file with markdown in it. An HTML error page served with a 200 status doesn’t.

What we couldn’t read

3 sites (Rippling, Personio and Axonius) answered every request with a security checkpoint. They’re excluded from the robots.txt and llms.txt figures.

5 more (Adobe, Gusto, Midjourney, Rubrik and ServiceNow) refused a script and served the same files to an ordinary browser, so we read them there.

9 homepages and 6 pricing URLs returned a bot wall or an empty shell. They’re excluded from the schema and pricing figures.

Limits of the method

This is one day and one host per company. It doesn’t cover docs or app subdomains, and it can’t see firewall rules.

Path-level rules don’t count as a block. Calendly, for one, limits CCBot to its pricing page and its llms.txt file, and we count that as allowed.

The dataset has one row per company and one column per measure. Download the full benchmark as a CSV and check any figure against it.

Frequently Asked Questions

Do SaaS companies block AI crawlers

Rarely. 9 of 161 leading SaaS companies fully block at least one AI crawler in robots.txt.

All 9 block training crawlers only. None blocks a search or user-fetch agent from OpenAI, Anthropic or Perplexity.

Should a SaaS company block GPTBot

Blocking GPTBot opts your pages out of OpenAI model training. It doesn’t remove you from ChatGPT search, which uses OAI-SearchBot.

3 of 161 companies in this sample block it. That’s a content-licensing decision, and it has no bearing on whether an assistant can cite you.

How many SaaS companies have an llms.txt file

102 of the 161 we could read, or 63.4%. Among public cloud companies the share is 70.8%.

Does robots.txt stop an AI crawler

It asks, and it doesn’t enforce. Crawlers that follow the standard obey it, and anything stricter needs a firewall or bot-management rule.

That also works in reverse. A firewall can block a crawler your robots.txt allows, so check your logs as well as your file.

What is a Content-Signal line in robots.txt

It’s a line that states how you want your content used: for search, for AI answers, or for AI training. Each is set to yes or no.

17 of 161 files in this sample carry one. It’s a stated preference, and a crawler can still ignore it.

Jordan Ellis
Written by

Jordan Ellis

Jordan Ellis is an AI search visibility specialist and content strategist with over 8 years of experience in B2B digital marketing. Focused on the intersection of content strategy and large language model optimization, Jordan writes about how brands can build lasting presence in AI-generated recommendations. Before specializing in AI visibility, Jordan led SEO and content programs for SaaS and FinTech companies across the US and Europe.

Leave a Reply

See where AI answers put your brand today.

Twenty minutes with a senior strategist: where the major engines point buyers in your category, who gets named instead of you, and a straight read on what a programme would change. No pitch deck.

Book a strategy call

A senior strategist replies within one business day.