Ask a SaaS marketing team whether they let AI crawlers in. Most can’t answer without opening a file nobody has touched in a year.
So we opened it for them. On 6 October 2026 we read the robots.txt of 164 leading cloud and SaaS companies.
9 of 161 block at least one AI crawler, and none of them block the crawlers behind AI search answers. Every block we found targets model training. 127 of the 161 files don’t mention AI crawlers at all.
The sample is the Forbes Cloud 100 list plus the 65 companies in the BVP Nasdaq Emerging Cloud Index. You can download every row at the end.
The Headline Numbers
Blocking is rare, and llms.txt is common. Here’s the full set.
| Signal | Companies | Share |
|---|---|---|
| robots.txt names an AI crawler | 34 of 161 | 21.1% |
| robots.txt blocks a training crawler | 9 of 161 | 5.6% |
| robots.txt blocks a search or user-fetch agent | 0 of 161 | 0% |
| Opts out of AI training by any method | 14 of 161 | 8.7% |
| Publishes an llms.txt file | 102 of 161 | 63.4% |
| Organization schema on the homepage | 118 of 155 | 76.1% |
| A page at /pricing | 107 of 158 | 67.7% |
| A price in the pricing page HTML | 66 of 158 | 41.8% |
The denominators differ because a few sites refused to be read. The method section lists each one.
What SaaS Robots.txt Files Say About AI Crawlers
Most say nothing. 127 of 161 files (78.9%) carry no rule for any of the 16 AI user agents we checked.

Silence means access. A crawler with no rule of its own follows the default group, and no company in the sample closes its whole site by default.
25 companies (15.5%) name AI crawlers only to let them in. The default already allows them, so the line changes nothing technically. It reads as a statement of intent.
Who blocks, and what
9 companies fully disallow at least one AI crawler: Canva, Figma, Forter, Midjourney, Notion, Postman, ServiceNow, Together AI and nCino.
3 block GPTBot (Canva, Midjourney and Figma). 2 block ClaudeBot. 3 block Google-Extended, Google’s opt-out token for Gemini.
The most-blocked crawler is Bytespider, ByteDance’s, with 6 blocks. CCBot, from Common Crawl, has 5.
Postman’s file gives its reason in a comment: “High volume, no citation, no traffic back.”
Training Is Blocked, Search Is Not
AI companies run separate crawlers for separate jobs. Each one can be controlled on its own line.
OpenAI’s crawler documentation lists GPTBot for training, OAI-SearchBot for search results, and ChatGPT-User for pages a person asks it to open. Anthropic’s crawler guidance and Perplexity’s crawler documentation describe the same split.

Not one of the 161 files fully blocks a search or user-fetch agent. OAI-SearchBot is named in 23 files, ChatGPT-User in 25 and PerplexityBot in 29. None of them shut it out.
The 9 blockers all drew the line in the same place. They refuse to be training data, and they still want to appear when a buyer asks an assistant for a tool.
For a software company that’s the sensible position. Blocking a search agent removes you from the answer and leaves your competitors in it.
The one company that rations search agents
Figma is the exception worth knowing. It doesn’t block AI search agents, and it doesn’t give them the run of the site.
Its robots.txt disallows everything for five search and user-fetch agents, then lists about 9,000 exact URLs they may read. That’s an allow-list, and it’s the only one in the sample.
Content-Signal lines say the same thing another way
17 files (10.6%) carry a Content-Signal line. It’s a robots.txt extension from Cloudflare’s Content Signals Policy that states a preference for three uses: search, AI input and AI training.
11 of them set ai-train=yes and 5 set ai-train=no. Forter sets both, for different crawler groups.
Add those to the robots.txt blocks and 14 of 161 companies (8.7%) opt out of AI training one way or the other.
llms.txt Adoption Among SaaS Leaders
102 of 161 companies (63.4%) publish an llms.txt file. That’s more than eleven times the number that block anything.
llms.txt is a proposed plain-text file listing a site’s most useful pages for language models, set out in the llms.txt proposal. The median file in our sample is about 13 KB.
Public companies are ahead. 46 of 65 index constituents have one (70.8%), against 57 of 97 Cloud 100 companies (58.8%).
Adoption isn’t proof of effect. Publishing the file shows intent, and it doesn’t show that any assistant reads it. Our guide covers what llms.txt is and what it can’t do.
We also logged who reads llms.txt on our own site.
What a Crawler Can Read on the Page
Access is only half of it. We also checked the raw HTML each server returns, which is what a crawler gets if it doesn’t run JavaScript.

118 of 155 readable homepages (76.1%) carry Organization schema. 111 (71.6%) include sameAs links to official profiles. Those two tell a machine who the company is.
Product-level markup is rarer. 27 homepages (17.4%) carry Product or SoftwareApplication schema, and 6 carry FAQPage.
Pricing is the gap
107 of 158 companies (67.7%) have a page at /pricing. Only 66 (41.8%) show a currency amount in that page’s HTML.
The other 41 pricing pages load the number with JavaScript or don’t publish one. A crawler reading the HTML finds no price on either kind.
When the vendor’s own page has no number, an answer about cost has to come from somewhere else.
What to Do With This
Check your own file before anything else. Access isn’t a way to get ahead, because nearly every leading SaaS company already allows AI search crawlers.
- Open your robots.txt and search for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot.
- Remove any rule that disallows a search or user-fetch agent. No company in this sample has one.
- Block the training crawlers by name if you want out of training, and leave the search agents alone.
- Load your pricing page with JavaScript off. A number that disappears is a number a crawler can’t read.
One caveat: robots.txt is a request, and a firewall rule can still block a crawler the file allows. 8 sites here refused our own script.
Your server logs show who gets through. We cover how to check which AI bots crawl your site.
Once the door is open, what decides whether you’re named is what other sites say about you. That’s the work behind GEO for SaaS brands.
How We Collected the Data
The sample is 164 companies: the 2025 Forbes Cloud 100 and the 65 constituents of the BVP Nasdaq Emerging Cloud Index on 6 October 2026. Netskope is on both lists and counted once.
For each company we requested four URLs on the host its homepage resolves to: /robots.txt, /llms.txt, the homepage and /pricing.
We parsed robots.txt by the group and longest-match rules in RFC 9309, the robots.txt standard. A crawler counts as blocked when its own group disallows the homepage and four other common paths.
The 16 user agents are PerplexityBot, ClaudeBot, ChatGPT-User, GPTBot, OAI-SearchBot, Google-Extended, CCBot, Applebot-Extended, Perplexity-User, Meta-ExternalAgent, Claude-SearchBot, anthropic-ai, Bytespider, cohere-ai, Claude-User and Amazonbot.
An llms.txt counts when the server returns a text file with markdown in it. An HTML error page served with a 200 status doesn’t.
What we couldn’t read
3 sites (Rippling, Personio and Axonius) answered every request with a security checkpoint. They’re excluded from the robots.txt and llms.txt figures.
5 more (Adobe, Gusto, Midjourney, Rubrik and ServiceNow) refused a script and served the same files to an ordinary browser, so we read them there.
9 homepages and 6 pricing URLs returned a bot wall or an empty shell. They’re excluded from the schema and pricing figures.
Limits of the method
This is one day and one host per company. It doesn’t cover docs or app subdomains, and it can’t see firewall rules.
Path-level rules don’t count as a block. Calendly, for one, limits CCBot to its pricing page and its llms.txt file, and we count that as allowed.
The dataset has one row per company and one column per measure. Download the full benchmark as a CSV and check any figure against it.
Frequently Asked Questions
Do SaaS companies block AI crawlers
Rarely. 9 of 161 leading SaaS companies fully block at least one AI crawler in robots.txt.
All 9 block training crawlers only. None blocks a search or user-fetch agent from OpenAI, Anthropic or Perplexity.
Should a SaaS company block GPTBot
Blocking GPTBot opts your pages out of OpenAI model training. It doesn’t remove you from ChatGPT search, which uses OAI-SearchBot.
3 of 161 companies in this sample block it. That’s a content-licensing decision, and it has no bearing on whether an assistant can cite you.
How many SaaS companies have an llms.txt file
102 of the 161 we could read, or 63.4%. Among public cloud companies the share is 70.8%.
Does robots.txt stop an AI crawler
It asks, and it doesn’t enforce. Crawlers that follow the standard obey it, and anything stricter needs a firewall or bot-management rule.
That also works in reverse. A firewall can block a crawler your robots.txt allows, so check your logs as well as your file.
What is a Content-Signal line in robots.txt
It’s a line that states how you want your content used: for search, for AI answers, or for AI training. Each is set to yes or no.
17 of 161 files in this sample carry one. It’s a stated preference, and a crawler can still ignore it.


