Resource

llms.txt: We Checked 3,254 SaaS Sites. Adoption Is Exploding. We Found No Evidence It Matters.

Original Research · LLM SEO · August 2026

Honest framing first: a lot of people are searching for llms.txt right now, and that is why this article exists. Most of what those searches return is either a sales page for a generator or an opinion with no data behind it. So we did what we keep doing with LLM SEO questions: we measured it. We crawled 3,254 B2B SaaS websites for llms.txt adoption, fingerprinted every file we found, checked each site’s robots.txt treatment of AI crawlers, and then cross-referenced adoption against real ChatGPT citation data from 120 SaaS buying queries.

The short version: SaaS is adopting llms.txt at roughly five times the rate of the general web, a meaningful slice of that adoption is accidental, and we found no relationship between having the file and being cited by AI. We publish one ourselves anyway. This article explains the spec, the data, and why “it probably does nothing, do it anyway” is the honest position.

Definition

llms.txt is a plain markdown file served at a website’s root (yoursite.com/llms.txt) that gives large language models a curated index of the site: an H1 with the site name, a one-line summary in a blockquote, and H2 sections of links with short descriptions. Proposed by Jeremy Howard of Answer.AI in September 2024, it is a voluntary convention, not a ratified standard, and no major AI lab has committed to reading it.

Key takeaways
  • 42.3% of the 3,254 SaaS sites we checked serve a valid llms.txt. The best published general-web figure is 8.7% of top-1,000 sites, so SaaS is adopting at roughly five times the web rate.
  • SEO plugins have made the file a checkbox: at least 17% of adopters run a plugin-generated llms.txt (Yoast, AIOSEO, Rank Math), the same route we use ourselves.
  • Companies with llms.txt were cited by ChatGPT at effectively the same rate as companies without one (61.5% vs 58.2%), and the tiny gap disappears once you control for company size.
  • The sites blocking AI crawlers are negotiating, not hiding: 95% of blockers restrict only training bots while leaving AI-search retrieval bots free, and AI-native startups block at eight times the enterprise rate.
  • Our verdict: no evidence it improves AI visibility today, but it costs ten minutes and has one genuine use case (developer documentation). Do it, then spend your real effort on authority.

What llms.txt actually is, and what it is not

The llmstxt.org proposal is short enough to read in five minutes, and most of the arguments about it come from people who have not. The file is a reading list, not a gate. robots.txt tells crawlers where they may not go. sitemap.xml lists every URL with no context. llms.txt sits between them: a human-readable, prioritised map of what matters on your site, written in the one format every language model parses natively, markdown.

A minimal valid file is an H1 with your company name, a blockquote summarising what you do, and a few H2 sections of links with one-line descriptions. There is an optional companion, llms-full.txt, which inlines the full text of your key pages into one document. Only 7% of our sample bothers with it.

What llms.txt is not: a ranking lever, an access control, or a Google requirement. Google’s John Mueller has said plainly that no major AI system currently uses it, comparing it to the meta keywords tag, and Gary Illyes confirmed in July 2025 that Google does not support it and is not planning to. Hold that thought, because Google’s own behaviour turns out to be more complicated.

Finding 1: SaaS is adopting llms.txt at five times the web rate

Of 3,254 B2B SaaS domains that responded to our crawler, 1,377 serve a valid llms.txt file. That is 42.3%, a number that surprised us enough to re-validate the crawler twice. Every positive required an HTTP 200 at the exact path, non-HTML content, and markdown structure, so single-page-app catch-all pages did not count.

llms.txt adoption by segment, August 2026Citation Gap Report cohort (n=146)62.3%Enterprise SaaS leaders (n=89)53.9%AI-native startups (n=127)47.2%Traffic-qualified mid-market (n=2,917)41.1%General web, Tranco top 1,000 (Rankability)8.7%
EMGI llms.txt adoption study, August 2026: 3,254 responding B2B SaaS domains, validated files only. General-web benchmark: Rankability.

For calibration: Rankability’s check of the Tranco top 1,000 websites found 8.7% adoption, and Originality.ai tracked around 36,000 adopting sites across the whole web as of late 2025. Our sample is deliberately not the whole web: it is traffic-qualified B2B SaaS spanning fifteen-plus categories, from HR tech, martech and fintech through devtools, legaltech, edtech, healthtech, data infrastructure, security and revenue operations, ranging from seed-stage AI-native startups to the enterprise leaders of each category. In other words, the most AI-aware commercial population on the internet, and the one whose buyers are already asking AI who to trust.

The one-line version, if you are quoting this: two in five B2B SaaS companies now publish an llms.txt file, roughly five times the adoption rate of the general web. And it climbs with company size: 54% among enterprise leaders against 41% across the broader mid-market pool.

Finding 2: a chunk of that adoption is accidental

Part of why SaaS adoption is so high: the major WordPress SEO plugins turned llms.txt into a checkbox. Yoast, AIOSEO and Rank Math can all generate the file automatically, which is exactly how we run ours at EMGI, and for most sites it is the sensible default. But it also means adoption statistics now mix two very different populations: teams that chose the file, and sites where a plugin shipped it.

We fingerprinted all 1,377 positive files for generator signatures. At least 239 of them, 17.4%, were machine-generated by an identifiable SEO plugin: Yoast (155), AIOSEO (43), Rank Math (37), plus a handful of other generators. Not everyone is watching what their plugin publishes, either: one file in our sample was still serving localhost URLs from the dev environment that generated it.

Who wrote the file? Generator fingerprints in 1,377 llms.txt filesNo generator signature found82.3%Yoast SEO11.3%AIOSEO3.1%Rank Math2.7%Other generators0.6%
EMGI llms.txt adoption study: signature scan of all 1,377 valid files. Unsigned generators are undetectable, so 17.4% machine-made is a floor.

Treat 17.4% as the floor, not the ceiling: plugins without signatures and agency-installed generators are invisible to fingerprinting. The practical point for anyone citing adoption statistics, including ours: a share of “adoption” is a plugin default rather than a decision, and no study can fully separate the two. Whichever camp a site is in, the fix is the same: know what the file says, because it describes your company to any machine that reads it.

Finding 3: the blockers are not anti-AI. They are negotiating.

Because we were fetching robots.txt anyway, we checked how each site treats eight AI crawlers, from GPTBot and ClaudeBot to ByteDance’s Bytespider. 141 sites, 4.3% of the sample, block at least one AI crawler outright. What they block, and what they conspicuously leave alone, tells you what SaaS companies actually believe about AI search.

AI crawlers blocked in robots.txt across 3,254 SaaS sitesBytespider (ByteDance)114CCBot (Common Crawl)99GPTBot (OpenAI)86Google-Extended72ClaudeBot (Anthropic)71PerplexityBot7
EMGI llms.txt adoption study: sites with an explicit Disallow for each crawler. 134 of 141 blockers restrict training bots only, leaving AI-search retrieval bots free.

Look at the pattern. 134 of the 141 blockers restrict only training-class crawlers, the bots that harvest content for model training, while leaving the retrieval bots that power live AI answers (PerplexityBot, OpenAI’s search crawler) completely free. Just 7 sites block retrieval. This is not confusion; it is a position: cite me, do not train on me. These companies want to appear in AI answers, with attribution, while withholding their content from the training corpora that make attribution unnecessary.

The sub-plots support the same reading. 28 sites block only Bytespider, ByteDance’s crawler, a pure trust judgement about one company. 65 run the full training lockout (GPTBot, ClaudeBot and Common Crawl together), the posture publishers adopted when content licensing became a market. And the most counterintuitive split in the whole study: AI-native startups block AI crawlers at 8.7%, eight times the rate of enterprise SaaS leaders at 1.1%. The companies building on these models are the most protective of what the models are allowed to eat. They know exactly where the value is.

Which also reframes the 74 sites that publish an llms.txt while blocking an AI crawler. A few are genuinely contradicting themselves, but for most it is the same negotiation stated twice: the llms.txt says “here is what to cite”, the robots.txt says “you may not keep it”. Whether the AI companies honour either document is, of course, a different question.

My own view, for what it is worth: I would not block any of them. Blocking training crawlers keeps your content out of the training data, but being in the training data is the opportunity. Retrieval citations are rented visibility; they last as long as your pages keep ranking. Training-data presence is owned: when the next generation of models ships, the brands woven through that corpus are the ones the model reaches for before it runs a single search. We saw this first-hand in our fan-out research, where Gemini answered most buying prompts purely from memory. If you are a larger company with a deep content library, an open door to training crawlers is a chance to influence how an entire category gets described for years, and blocking it trades that away to make a licensing point. Your content is getting retrieved either way; you may as well be in the model’s memory too.

The one-line version: SaaS companies that block AI are blocking the training bots, not the answer engines.

Finding 4: we found no relationship between llms.txt and AI citations

Here is the test everyone actually cares about. In April we captured which of 150 SaaS companies ChatGPT mentions across 120 real buying queries for the SaaS AI Citation Gap Report. Those same companies are in this crawl, which lets us ask directly: do the ones with llms.txt get cited more?

No. Companies with the file were mentioned in at least one answer 61.5% of the time; companies without, 58.2%. Average mentions per company: 4.7 with, 4.6 without. And even that sliver evaporates under a size control: dominant brands get cited 100% of the time with or without the file, and mid-tier adopters actually averaged slightly fewer mentions than mid-tier non-adopters. Whatever drives AI citations, it is not this file.

Does llms.txt correlate with ChatGPT citations? (146 SaaS companies)With llms.txt: cited in at least one answer61.5%Without llms.txt: cited in at least one answer58.2%With llms.txt: avg mentions per company (x10)47%Without llms.txt: avg mentions per company (x10)46%
Cross-reference of this crawl against the Citation Gap Report ChatGPT capture (120 buying queries, April 2026). Averages shown x10 for scale: 4.7 vs 4.6 mentions.

The subgroup detail is worth a moment, because it shows what noise looks like. Among challenger brands, adopters were cited more often than non-adopters (26% vs 14%), which an advocate would headline. But among mid-tier brands the direction reverses: adopters averaged fewer mentions than non-adopters (3.7 vs 4.4). When an effect flips sign between adjacent subgroups on samples this size, you are looking at scatter, not signal. A real lever pulls in one direction.

There is also a mechanical reason to expect exactly this null. When an AI engine answers a buying question, it fans the question out into sub-queries and retrieves the pages that rank for them, a pipeline we mapped in our query fan-out capture. Nothing in that pipeline consults a root-level index file. The engine finds you through ranked, retrievable pages and the brand associations already in the model, which is why authority and mentions predict citations and configuration files do not.

This matches the strongest independent evidence available. Ahrefs analysed server logs across 137,210 domains in June 2026 and found 97% of llms.txt files received zero traffic, with AI retrieval bots accounting for about 1% of what little did arrive. The file is not being read at scale, so it is in no position to change outcomes at scale. Our correlation data and their log data say the same thing from opposite ends.

And step back for the simplest argument of all, no crawl required: a ten-minute file that 42% of your competitors already serve cannot be a source of significant advantage. Cheap, ubiquitous tactics never are. If llms.txt moved citations, the 1,377 adopters in our sample would be visibly outperforming the rest. They are not.

The honest caveats: our capture is one engine (ChatGPT), one moment in time, and 146 usable companies is a modest sample for subgroup analysis. Adopters also skew toward savvier marketing teams, which would bias the result in llms.txt’s favour if anything, making the null finding more telling, not less.

So why does anyone bother? The one real use case

Because in one corner of the internet, llms.txt genuinely works today: developer documentation. Coding assistants and agent tools do fetch these files from docs sites, which is why Mintlify generates them for every docs site it hosts and why Zapier, Cloudflare and Anthropic’s developer platform all maintain them. If your product has an API and your buyers include developers pasting your docs into Cursor or Claude, an llms.txt over your documentation is doing real work right now.

There is also a hedge argument, and a wrinkle that keeps it interesting: Google says it ignores llms.txt, yet Google’s own Chrome team published one, and Cyrus Shepard’s testing showed Google can read and even cite the files’ contents. Nobody serious claims it helps today. The claim is that the cost is ten minutes, the downside is zero, and conventions occasionally become standards.

How to create an llms.txt in ten minutes

Since you will probably do it anyway, here is the ten-minute version. An SEO plugin will do a perfectly good job of this automatically, and honestly it is never that deep; the only part genuinely worth doing by hand is the summary line.

1. Open a text file. First line: # Your Company Name. Second: a blockquote (>) with one honest sentence on what you do and for whom. This summary is the highest-leverage line in the file; write it like positioning, not a slogan.

2. Add H2 sections for your key content groups (Products, Docs, Pricing, Research), each containing markdown links with a one-line description per page. Ten to thirty links. Curate; do not dump your sitemap.

3. Point links at your most citable pages: pricing, comparisons, original research, documentation. The pages an AI could quote, not the pages you wish people visited.

4. Serve it at yoursite.com/llms.txt as plain text. Or take the genuinely easiest route: if you run WordPress, Yoast, AIOSEO and Rank Math can all generate and serve the file automatically, and letting the plugin handle hosting while you edit the content is the best of both. Either way, audit the summary line personally, because that sentence is how your company describes itself to any machine that reads the file.

5. Skip llms-full.txt unless you run developer docs. Revisit twice a year.

The deeper play, and the one our data supports, is what we call semantic brand positioning: making sure every surface that describes your company, from your llms.txt summary to your directory listings to your comparison pages, says the same thing in the same language. The file is one small surface of that. Our directory study found directories plus topical authority produce five times more ChatGPT citations than authority alone; consistency across surfaces is what makes the machine confident about what you are.

The key numbers, stated plainly

If you are citing this study, these are the headline statistics in one place (EMGI llms.txt adoption study, August 2026, n = 3,254 B2B SaaS domains):

  • 42.3% of B2B SaaS websites serve a valid llms.txt file, roughly five times the best published general-web adoption rate (8.7% of top-1,000 sites, Rankability).
  • Adoption rises with company size: 53.9% among enterprise SaaS leaders vs 41.1% across the broader traffic-qualified mid-market.
  • At least 17.4% of llms.txt files were auto-generated by SEO plugins (Yoast, AIOSEO, Rank Math), not deliberately created.
  • Only 7% of SaaS sites serve the llms-full.txt companion file.
  • 95% of SaaS sites that block AI crawlers (134 of 141) block only training-class bots while leaving AI-search retrieval bots free: a deliberate “cite me, do not train on me” position.
  • AI-native startups block AI crawlers at 8.7%, eight times the 1.1% rate of enterprise SaaS leaders.
  • Companies with llms.txt were cited by ChatGPT at the same rate as companies without: 61.5% vs 58.2%, a gap that disappears when controlling for company size.
  • The most-blocked AI crawler among SaaS sites is ByteDance’s Bytespider, ahead of CCBot and GPTBot.

Frequently asked questions

Is llms.txt a real standard?

It is a real, widely adopted convention but not a ratified standard. Jeremy Howard of Answer.AI proposed it in September 2024, the spec lives at llmstxt.org, and no standards body or major AI lab has formally committed to it. Adoption is real (42% of the SaaS sites we checked), commitment from the AI companies is not.

Does Google use llms.txt?

Google Search says no, twice on the record: John Mueller compared it to the meta keywords tag, and Gary Illyes said in July 2025 that Google does not support it and is not planning to. At the same time, Google’s Chrome team published its own llms.txt page and independent testing shows Google can read and surface the files’ contents. The accurate summary: not a ranking input, occasionally read anyway.

Is llms.txt worth doing?

As a ten-minute task, yes; as a strategy, no. Think about it structurally: if a free file that takes ten minutes and that 42% of your competitors already serve could deliver a significant visibility advantage, that advantage would already be gone. Our data agrees: no relationship with ChatGPT citations (61.5% vs 58.2% cited, converging to zero under a size control), and Ahrefs’ server logs show 97% of these files receive no traffic at all. Do it for the hedge and the docs use case, then invest your actual effort in authority and citable content, where scarcity still exists.

What should an llms.txt file contain?

An H1 with your company name, a blockquote with a one-sentence description of what you do, and H2 sections of curated markdown links with short descriptions: pricing, comparisons, documentation, research. Ten to thirty links beats a sitemap dump, and the summary sentence deserves the most care, since it is how machines will describe you.

Do I already have an llms.txt without knowing it?

Quite possibly. Yoast, AIOSEO and Rank Math can all generate one automatically, and they accounted for at least 17.4% of the files in our study. Check yoursite.com/llms.txt now; if a file exists that you never wrote, audit it, because one file in our sample was still linking to its developer’s localhost.

Methodology and limitations

We crawled 3,325 B2B SaaS domains in August 2026: a traffic-qualified prospecting database of 2,979 SaaS companies, the 150 companies from our Citation Gap Report, and a stratified supplement of 300 enterprise and AI-native SaaS firms. 3,254 responded. A valid llms.txt required HTTP 200 at the exact path, non-HTML content and markdown structure. All positives were re-fetched and fingerprinted for generator signatures. robots.txt rules were parsed for eight AI crawlers. Citation cross-referencing used the ChatGPT capture behind our Citation Gap Report (150 companies, 120 buying queries, April 2026). Limitations: one crawl date, apex domains only (docs subdomains not checked, which understates docs-driven adoption), one AI engine in the citation test, and generator fingerprinting only catches signed files. The full aggregate results are available on request; we do not publish the raw domain list.

The wider programme this belongs to: we run original captures on how AI engines actually behave, from the query fan-out captures behind our AI-search research to which subreddits Google’s AI quotes. The consistent lesson across all of them is that visibility follows authority and retrievable content, not configuration files. If you want the same lens on your own brand, that is the work we do.

Matt Emgi is the founder of EMGI Group, a SaaS link building and AI visibility agency. He publishes original research on how AI engines select and cite sources, including the SaaS AI Citation Gap Report, the Reddit Citation Study, and the query fan-out capture series.
Work with EMGI

We build the authority that gets SaaS brands cited: editorial links, authority content, and brand visibility across the surfaces AI engines trust. Book a call with Matt and we will show you where you stand today.