WritingContext engineering6 October 20269 min

llms.txt does nothing.
That was always the wrong layer.

The market measured a signpost and concluded the road doesn't exist. The manifest was never the mechanism. Whether your text is actually there — structured, and attributable to its source — is.

Brian HumeFractional CTO & AI Systems Architect

408
requests for llms.txt
of
515,382,577
AI bot requests, 90 days

Limy.ai monitoring across the brands it tracks, 2026. The crawlers in that traffic — GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended — overwhelmingly fetched HTML.

Eighteen months after llms.txt was proposed as the way to make a website legible to language models, the measurements are in, and they are unkind.

Roughly one site in ten has the file — a SE Ranking study of 300,000 domains puts adoption at 10.13%. Of the 50 domains most often cited by AI answer engines, one has it. HTTP Archive's Web Almanac found that around 40% of the llms.txt files that do exist are plugin-generated defaults nobody wrote. And in the traffic figure above, from a monitoring set of over half a billion AI bot requests, 408 asked for the file directly.

Google said on the record, in July 2025, that it does not support llms.txt and is not planning to. No major provider — OpenAI, Anthropic, Google, Meta, Mistral — has publicly committed to reading it in a production answer surface.

So the verdict forming across the SEO industry is: llms.txt is dead, generative engine optimisation was a fad, back to work.

The data is right. The conclusion is wrong. And the reason it is wrong is worth understanding, because the same mistake is about to be made with structured data, and then with whatever comes after that.

What llms.txt actually is

Strip the discourse away and llms.txt is a table of contents. A Markdown file at the root of a site that says: here is what this site is about, here are the pages worth reading, here is the order to read them in.

That is useful. It is also the smallest possible contribution to whether a model can use your content. Nobody has ever claimed that a sitemap makes a page rank. A sitemap tells a crawler where to look. What it finds when it looks is a separate question, and it is the only question that matters.

llms.txt
A plain-text manifest at a site's root listing the pages a language model should read, with a short description of each. A pointer, not a payload.
Manifest
Any file whose job is to say what exists and where. Sitemaps, robots.txt and llms.txt are manifests. None of them contains the content they describe.
Mechanism
The thing that actually determines whether a model can retrieve, understand and attribute a passage. It lives in the page, not in the file that points to it.

So when the measurements show that crawlers ignore the manifest and fetch the HTML instead, that is not evidence that machine-legibility is a myth. It is evidence that the crawlers are doing exactly what you would do: skipping the index and reading the book.

People measured a signpost and concluded the road doesn't exist.

The layer everyone skips

Here is the precondition that almost no discussion of AI visibility mentions, and it is the one that decides everything downstream.

If your text is not in the HTML the server sends, a crawler cannot read it.

A large share of the modern web ships an empty <div id="root"> and a JavaScript bundle, and renders the words in the browser after the fact. A human never notices. A retrieval system fetching the page at citation time frequently does: it reads the source, finds a shell, and moves on. Whatever your llms.txt promised, the page it pointed to was blank on arrival.

My own site, brianhume.dev, is a React application. It is also prerendered at build time, so the full text of every page travels inside the source HTML and does not depend on JavaScript executing anywhere. That is not a performance optimisation — it is the condition for being readable by anything that is not a browser. It costs one build step. Most sites that have an llms.txt have not taken it.

The same applies to the structured-data advice currently replacing the llms.txt advice. JSON-LD is valuable. But a retrieval system selecting a passage to cite is reading visible HTML at fetch time; markup describing a page whose text is not there describes nothing. You cannot label your way out of an empty page.

Three layers, one source

The architecture I run on brianhume.dev and on Mr. Wolf, a learning platform built to serve human readers and a language model from the same content, has three layers. llms.txt sits in the middle one, as one line item.

1 · SOURCE One Markdown file written once, consumed twice definition extraction 2 · HUMAN LAYER The page people read semantic HTML, prerendered in source JSON-LD per entity type llms.txt + sitemap: what to read, in order real headings, lists and tables 3 · MODEL LAYER The same text, retrievable chunked by definition, not by length each chunk keeps its source URL per-call trace: which chunk entered which generated answer citation reconstructible back to source
The manifest is one line in the middle layer. The mechanism is the bottom one: if every retrieved chunk carries the URL it came from, you can prove where an answer originated instead of assuming.

Layer one: a single source

The content is written once, in Markdown. Before it is served anywhere, a definition-extraction pass delimits each concept — this term, this meaning, these boundaries — so that a definition is a unit rather than a run of sentences a chunker might split in half.

Layer two: the human page

Semantic HTML with real headings, lists and tables rather than visual layout; prerendered so the text is in the source; JSON-LD per entity type so a machine knows what kind of thing it is looking at; and, yes, an llms.txt and a sitemap saying what to read and in what order. On brianhume.dev the robots file explicitly admits GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, Applebot-Extended and CCBot. Generative crawlers are a first-class audience, treated as one.

Layer three: the same text, retrievable

This is the point of the architecture, and it is where almost everyone stops short. The same content is chunked by definition rather than by character count, each chunk keeps its source URL, and every model call records which chunks entered which generated answer.

If each fragment knows where it came from, you can demonstrate where a generated answer originated instead of guessing. It is the same discipline I apply to AI systems in production — every model call traced with its cost, its prompt version and its evidence — pointed at content instead of at a pipeline.

A summary of a summary cannot be audited. A chunk with a source URL can.

The objection worth taking seriously

"Fine, but is there any evidence that any of this produces citations?"

Honest answer: I am not going to show you a traffic chart, because I do not yet have one I would stand behind. What I have is a working implementation and a set of screenshots proving it exists — in Google's AI mode, in rich-results testing, in the context payloads a model actually receives. That is evidence the architecture is real. It is not evidence of lift, and I am not going to dress it up as one.

Which is more than the llms.txt verdict can say. The studies that killed the manifest measured whether crawlers request the file. None of them measured whether a prerendered, definition-chunked, source-attributed page is cited more often than a client-rendered one with a JSON-LD label on it. That is the experiment that would settle something, and as far as I can find, nobody has run it.

Until someone does, the honest position is architectural, not statistical: build the thing that makes citation possible, and do not mistake the manifest for the thing.

One job, two surfaces

Here is the part that should make this less controversial than it sounds.

Every property that makes a page usable by a language model — text served in the source, semantic hierarchy, delimited definitions, explicit entities — is the same property that makes it usable by a search crawler. These were best practices before anyone said "generative engine". They were best practices before anyone said "search engine optimisation", if you go back far enough; they are what a well-made document looks like.

So this is not a new discipline competing with the old one for budget. It is one piece of work, measured on two surfaces. If you have built your site so a model can read it, you have built it so Google can read it, and vice versa. The teams that treated llms.txt as a separate initiative — a file to add, a box to tick — got the result you would expect from a box ticked. The ones that treated it as one line in a structural decision got a site that works on both surfaces, and will keep working when the next manifest format is proposed and then declared dead.

llms.txt is on my site. It will stay there. It costs nothing and it is a reasonable courtesy to a crawler that reads it. But if it vanished tomorrow, nothing about how my content is retrieved would change — because it was never the layer doing the work.

Stop optimising the metadata. Make sure the text is there first.

Sources

  1. Limy.ai, LLMs.txt in 2026: The Full Guide — 515,382,577 LLM bot traffic events over a 90-day window; 408 direct requests for llms.txt; summary of provider positions.
  2. SE Ranking, llms.txt adoption study — 10.13% adoption across 300,000 domains; 1 of the 50 most AI-cited domains has the file.
  3. HTTP Archive, 2025 Web Almanac — approximately 40% of published llms.txt files are plugin-generated defaults.
  4. Gary Illyes and John Mueller, Google, July 2025 — Google does not support llms.txt and is not planning to.