---
title: "Does llms.txt get you cited by ChatGPT? What Google and OpenAI actually say"
url: https://samambai.com/blog/does-llms-txt-work/
date: 2026-10-03
author: samambai
---

llms.txt does not get a business cited by ChatGPT, Gemini, or Google's AI. Google says Search ignores the file. OpenAI tells sites to allow OAI-SearchBot and the published IP ranges. The file still works as an index for coding agents and docs.

No. Publishing an `llms.txt` file does not get a business cited by ChatGPT, Gemini, or Google's AI. Google says Search does not use the file, and that keeping one neither helps nor hurts visibility or rankings [1] [2]. On the page where OpenAI explains how a site shows up in ChatGPT search, the documented step is different: allow OAI-SearchBot and allow the published IP ranges. That page does not list `llms.txt` as a condition for a citation [4]. The file is still useful in a narrow case, a coding agent that needs a map of the docs. That is why samambai keeps one.

A vendor that sells the file as a citation shortcut is selling an index. An index is not an answer, not a third-party source, and not a rank.

## What the file actually is

`llms.txt` is a proposal, not a required search standard. Jeremy Howard published the first version on 3 September 2024 and updated it to v2 on 10 August 2026 [3]. The point is to hand an agent a short Markdown map instead of making it parse a full HTML page, with the nav, the ads, and the scripts.

The spec is small. The only required section is an H1 with the name of the project or the site. After that you can add a blockquote summary, free text that does not open new headings, and H2 sections whose bodies are lists of links. Each list item is a link, then an optional note after a colon. A section titled Optional is the convention for links an agent may skip when the context window is tight [3].

The file can sit at the root (`/llms.txt`) or on any path, and it covers the URLs under that path. If two files apply, the spec says to use the more specific one. It sits next to `robots.txt` and a sitemap. It does not replace them. `robots.txt` says what a robot may fetch. A sitemap lists pages for indexing. `llms.txt` is read on demand, when an agent needs a topic while helping a person. Howard's expectation is inference, not training, though a training run could use the text too [3].

The proposal is plain about where the file is used most: software documentation, where a coding agent follows links to an API reference and a tutorial. The same shape can, in his words, describe a business, a personal site, or a school. That is a possible use. It is not a promise that ChatGPT will cite the business [3].

The proposal also asks for a clean Markdown twin of important pages, at the same URL with `.md` appended, plus `rel="alternate"` links or an HTTP `Link` header. Chrome Lighthouse, the spec says, audits for the file as part of its agentic browsing checks [3]. A developer-tool audit is not a Google ranking factor. Google's own guide, in the mythbusting section, tells site owners to ignore AI text files [1].

## What Google has said

"Optimizing your website for generative AI features on Google Search" was last updated on 10 July 2026. The mythbusting section says you do not need new machine-readable files, AI text files, markup, or Markdown to appear in Google Search, including its generative AI features, because Google Search itself does not use them. Discovering, crawling, and indexing a file does not mean the file gets special treatment [1].

The next sentences are the ones that matter if you are afraid to delete the file. It is fine to create and maintain an `llms.txt` (or a similar file) for other services or systems that use it. Doing so will neither harm nor help your visibility or rankings in Google Search, because Google Search ignores those files [1].

That note was added on 15 June 2026. The Search Central changelog records the reason in the same terms: the files are not needed for Google Search, and they will not negatively or positively affect visibility or rankings. Keeping them for other services or systems is fine [2]. This is not a paraphrase from a conference hallway. It is the official documentation changelog.

The same guide tells you to ignore, in the same breath, chopping pages into tiny pieces ("chunking"), rewriting copy only for AI systems, chasing inauthentic mentions, and treating a special schema type as a requirement for generative search. Structured data remains useful for classic rich results. There is no special schema.org markup for AI Overviews or AI Mode [1]. The full citation model, index then page then third-party trust then freshness, is in [what gets a business cited by AI](/blog/what-gets-a-business-cited-by-ai/).

On 15 May 2026 the same changelog clarified that existing spam policies also apply to generative AI responses in Google Search [2]. Stuffing hidden instructions into `llms.txt`, or publishing dozens of pages only to cover query variants, is not "file optimization." It sits under the spam rules Google says already cover the generated answer. The guide warns that separate pages for every variation of a query, built to manipulate rankings or generative responses, violate the scaled content abuse policy [1].

## What a log study actually measured

An official statement answers "does Google use this?" A log answers "did anyone request the file?" Those are different questions. The second one has a vendor study, not a Google measurement.

On 15 June 2026, Louise Linehan and Xibeijia Guan at Ahrefs published a read of Ahrefs Web Analytics and Bot Analytics [6]. The population was every domain in that product with traffic in May 2026: 137,210 domains. Ahrefs says those customers skew more technical and more SEO-aware than the web at large, so the adoption rate is an upper bound, not an internet average.

They looked for `/llms.txt` at the root returning HTTP 200, checked that the body was Markdown rather than an HTML error page, and counted requests to that path. They did not check whether each file matched the spec. The study measures the index file, and only the index file [6].

| What they measured | Figure in the Ahrefs study | How to read it |
| --- | --- | --- |
| Domains with traffic in May 2026 | 137,210 | Vendor base, skewed toward technical customers |
| Publish a valid `llms.txt` | 28% (38,360) | A ceiling, not a web average |
| Files with zero requests in May 2026 | 97% | No bot, no person |
| Requests that did arrive and came from bots | 96% | The other 4% were people |
| Requests from named AI tools, among the ones that arrived | 19.5% | GPTBot first, Claude-Code second |
| AI search retrieval bots (OAI-SearchBot, PerplexityBot, Claude search) | 1.1% of requests | The bots that cite barely read the index |
| Requests from people studying the file (GEO tools, scanners, research) | about 12% | Audit traffic, not citations |
| AI-bot requests to an `/llms.txt` that 404s | zero | Nothing goes looking for a file you did not publish |

Ahrefs' chart marks 38,360 domains with a valid file, 28% of the 137,210 base. The complement drawn on that same chart, 98,640, adds up to 137,000 with the 38,360. That is the chart rounding, not a second count. The roughly 3% that received any request hold the measured traffic, on the order of 1,100 domains and about 22,000 requests. Inside that small pool, Chrome's Lighthouse `llms.txt` audit accounted for about 1 in 1,000 fetches (22 requests) [6].

Two details should change a buying decision. First, "fetched" is not "used." Ahrefs writes that many bots may have downloaded the file without acting on it, so every AI percentage is a ceiling [6]. Second, Slack's link-preview bot fetched the file more often than PerplexityBot. A chat unfurl is not a citation.

Ahrefs also splits the readers. Agents and agent infrastructure were 10.5% of requests. Training crawlers were 5.3%, with GPTBot alone at 4.51% and ClaudeBot at 0.8%. User-triggered assistants were 2.5%. Retrieval bots were 1.1%, and inside that group OAI-SearchBot led at 0.74%. Claude-Code, Anthropic's coding agent, out-fetched every retrieval bot, every assistant, and every training crawler except GPTBot and an indexer named statespace-indexer, whose operator Ahrefs names without confirming the IP ranges [6].

Googlebot showed up about 900 times in May. Ahrefs reads that the way the guide does: Googlebot fetches URLs it discovers, the same way it fetches a sitemap. Those fetches do not show special interest in `llms.txt`, and the study cannot see whether any of that text later feeds Gemini [6]. That matches the guide. Indexing one more file type does not create special treatment [1].

On paths that returned 404, the split flipped. A valid file drew 96% bot traffic. A missing file drew 98% human traffic, and the AI-bot share of those 404s was zero. The person typing the URL into a browser, usually to check a competitor, is a person. An AI system does not hunt for an `llms.txt` you never published [6].

One security finding is worth keeping, without the drama. Among research bots (2.7% of requests), the largest identified itself as `prompt-injection-survey/1.0`. Someone is scanning the file as a prompt-injection surface, because agents are built to trust what the index says [6]. Reason enough to treat the file as public, short text. Not a reason to hide instructions in it.

A vendor study measures that vendor's logs. It does not measure how often ChatGPT cited a brand. It does not prove that deleting the file drops citations, or that publishing one raises them. It proves something simpler. In May 2026, almost nobody requested the file, and the requests that existed were rarely from the search bot that cites.

## What OpenAI actually documents

"Overview of OpenAI crawlers," as read on 2 October 2026, separates four agents. Each control is independent [4].

| Agent | What OpenAI says it is for | What it is not |
| --- | --- | --- |
| OAI-SearchBot | Surfaces sites in ChatGPT search results. A site that opts out is not shown in search answers, though it can still appear as a navigational link. | Not the `llms.txt` file |
| GPTBot | Crawls content that may be used to train foundation models. Disallowing it says the content should not train those models. | Does not decide search |
| ChatGPT-User | A visit triggered by a person in ChatGPT or a custom GPT. Because the person started the action, robots.txt rules may not apply. | OpenAI says it does not decide Search inclusion |
| OAI-AdsBot | Visits only pages submitted as ads, to check the page and the ad's relevance. The data is not used to train foundation models. | Not organic inclusion |

The written recommendation for search is specific: allow OAI-SearchBot in `robots.txt`, and allow requests from the IP ranges published at `https://openai.com/searchbot.json` [4]. GPTBot, ChatGPT-User, and OAI-AdsBot have their own lists. If a site allows both OAI-SearchBot and GPTBot, OpenAI may use a single crawl for both jobs so it does not crawl twice. A robots.txt change takes about 24 hours to propagate in their systems [4].

The sample OAI-SearchBot user agent on that page ends in `OAI-SearchBot/1.4` and points at `https://openai.com/searchbot`. When they fetch `robots.txt`, they may add a `robots.txt` marker to the user agent so logs can tell that fetch from the others [4]. The version number can change. The page says so.

Nothing on that page says "publish an llms.txt." The only mention of the filename is in the docs header, and the meaning is different: "For the complete documentation index, see llms.txt" [4]. That is an index of developer documentation. The file at `https://developers.openai.com/llms.txt`, fetched the same day, confirms it. The H1 is "OpenAI Developers." The blockquote describes the hub for the API, Ads, Codex, and the rest. The H2 sections point at other product `llms.txt` files and at `.md` guides: quickstart, reference, evals, agentic commerce [5]. That is the spec's format, used so an agent can find the right docs page. It is not a file telling ChatGPT to cite OpenAI when someone asks for a vendor.

Mixing up the four robots is the expensive mistake. Blocking GPTBot is not the same as leaving ChatGPT search. Allowing GPTBot does not put the page in the answer. ChatGPT-User may visit a URL a person asked for, and that visit is not `llms.txt` being "read by ChatGPT." Ads have their own robot and, in the docs, do not train foundation models. Showing up in a search answer follows the OAI-SearchBot line and the IP allowlist, and even then a citation is not guaranteed. The page still has to be retrieved, and the model still has to use it. The index file is not in that sentence.

## When the file is still worth having

It is worth having when someone will point an agent at the site and that agent needs a map. The spec describes the case: library docs, an API reference, a tutorial [3]. OpenAI does exactly that for developer documentation [4] [5]. The proposal says Anthropic's and Gemini's docs do the same, and it prints those URLs [3]. We did not fetch those two files for this article. What we did fetch is OpenAI's file and the spec.

It is also worth having as an internal index. A coding agent maintaining the site, or an agent inside the operation, spends less context if it starts from a short list of Markdown pages than if it starts from the HTML. The spec says to test the file by asking an agent questions with only that `llms.txt` as the starting point [3]. That test tells you whether the map is clear. It does not tell you whether you will be cited.

It is not worth having as a visibility project. The vendor base rate is blunt: 97% of existing files had no reader of any kind for a full month, and the bots that retrieve answers were 1.1% of the requests that existed [6]. Publishing the file does not put the domain on a radar. The 404s show the opposite [6].

### Why samambai keeps one

The studio publishes `https://samambai.com/llms.txt`. The file is a short map: what the studio does, public prices in US dollars, how an engagement runs, how to get in touch, and links to Markdown versions of the key pages. The reason is the same one behind OpenAI's developer docs. An agent that has already been sent to the site can read the service, the timeline, and the price without guessing at the HTML. The studio's own operation is AI-first: people plus agents for building, testing, and an independent audit, always with a human review. That internal agent is a plausible reader. So is a buyer who pastes the URL into a coding agent.

The file is not there to win a citation. Google says keeping it for other systems does not change Search [1] [2]. The vendor study says that when a reader shows up, it looks more like a coding agent than like a search bot [6]. Treating our file as proof of GEO would repeat the myth this page takes apart.

If the map goes stale, it lies to the only reader that matters. Price, timeline, and scope in `llms.txt` have to match the site. The spec asks for short language and a note on each link [3]. A stale index is worse than no index, because the agent that trusts it repeats the error. If the site says US$1,200 for a sprint and the file still says US$800, the agent quotes US$800.

## What to do instead

Citation follows a different path. The long version is [what gets a business cited by ChatGPT, Gemini, and Google's AI](/blog/what-gets-a-business-cited-by-ai/). The operational order, the order in which the work actually stalls, is short.

1. Access to the index of the engine you care about. On Google, the page has to be indexable and eligible for a snippet, and the site has to be included in Search Console's generative AI control. The guide says meeting the requirements does not guarantee indexing or serving [1]. On ChatGPT, search asks you to allow OAI-SearchBot and to allow the IP ranges in `searchbot.json` at the firewall or the CDN [4]. Blocking Googlebot "so they cannot train on us" and then expecting a citation in Google's AI contradicts the guide: Google's generative search is grounded in the Search index [1].
2. A page that answers the buying question. Price, timeline, what is included, and the date that fact is true. The guide wants non-commodity content: something you know because you did the work, not a summary of what every other page already says [1]. "AI solutions for your business" does not answer "what does a two-week automation sprint cost."
3. Trust that does not come from your own site. Bought mentions, a network of satellite domains, and hidden text in a "summarize with AI" button sit on the wrong side of the policy Google says covers the generated answer, clarified on 15 May 2026 [2]. The guide says chasing inauthentic mentions does not help, because core ranking and spam systems still apply to the generative feature [1].
4. Real freshness, not a fresh stamp. If the price changed, the page changes. Putting today's date on last year's copy is not an update.

Measure without fooling yourself. Track presence across many runs of the same question, and impressions in Search Console's generative AI report. The guide points at that report. It shows discovery, not a rank [1]. One screenshot is an anecdote. A third-party tool that promises internal Google metrics does not have them. The guide says so outright [1].

GEO at samambai is that work: find out what is already in the index, make the page answer, earn real corroboration, and keep a rhythm of updates. It is not an `llms.txt` install. Public prices are on [pricing](/pricing/): a Sprint at US$1,200 per two weeks, cancel whenever you want; a Project from US$3,000 with scope, timeline, and price fixed up front; a Partnership from US$900 a month. The first call is 30 minutes and creates no commitment. GEO for a specific site is quoted after that call, because it depends on what is already indexed. The service page is [GEO](/geo/).

A buyer in New York asking which integrator to hire, a buyer in Madrid comparing WhatsApp agents, and a buyer in Tel Aviv checking a fixed project price all hit the same wall. If the page does not state the price and the date, the index has nothing solid to repeat. The file at the root does not fill that gap.

## Mistakes that waste a week

| Mistake | Why it fails |
| --- | --- |
| Publish the file and expect a citation within a week | 97% of files in the study got no request in a month. Search bots were 1.1% of what did arrive [6]. |
| Assume Google rewards or punishes the file | The guide and the 15 June 2026 changelog say the opposite: ignored, no help, no harm [1] [2]. |
| Block GPTBot and conclude you have left ChatGPT | The controls are independent. Search is OAI-SearchBot. Training is GPTBot [4]. |
| Allow only ChatGPT-User | OpenAI says that agent does not decide Search inclusion [4]. |
| Generate the file from a plugin and never touch it again | The plausible reader is an agent that trusts the map. A stale map repeats the wrong price. |
| Count a fetch in the log as a win | Ahrefs says a fetch is not use [6]. |
| One URL per query variant, then list them all in the file | The guide treats volume built to manipulate the answer as scaled content abuse [1]. |
| Hide an instruction that tells the model to store a memory | The file is public. The study already saw a crawler identifying itself as prompt-injection research [6]. |

## What not to do

Do not swap the page that answers for an index file. Do not buy mentions. Do not stand up a domain network whose only job is to "cite" the brand. Do not promise, and do not hire anyone who promises, the first slot in ChatGPT. Do not use `llms.txt` as a place to instruct the model. Do not treat schema, chunking, or a cosmetic rewrite as a substitute for the index and for proof [1].

Do not delete the file because you are afraid Google will punish it. There is no documented penalty, and no documented prize [1] [2]. Delete it, or skip creating it, when nobody will point an agent at it and the cost of keeping it true is real. Fifteen lines that state the wrong price cost more than zero lines.

Do not read the Ahrefs study as if it were Google. It is one vendor's logs, on a technical customer base, for one month (May 2026). It is enough to retire the idea that "everyone is already reading llms.txt." It is not a citation rate.

## When not to hire samambai

Do not hire the studio if the ask is one of these:

- A guarantee of first place in ChatGPT, Gemini, or AI Mode. Nobody reading the docs can offer that. Google's guide says indexing does not guarantee serving [1].
- Paid media. Ad management is not part of what samambai sells. OAI-AdsBot is OpenAI's product for people who advertise in ChatGPT, and it is not work this studio sells [4].
- A site network, a mention package, or a "summarize with AI" button that tries to write memory into the assistant. That is the kind of manipulation Google's spam policy reaches inside the generated answer [2].
- A project whose deliverable is the `llms.txt` file, a special schema type, or a page chopped into chunks. Google says to ignore all three [1].
- A result in a few days on a site that is not in the index and has no third-party corroboration. Without the first gate, the file does not create the second.
- A new URL for every way of asking the same question. The guide calls that a weak strategy and, when the goal is manipulation, scaled content abuse [1].

Write to [contato@samambai.com](mailto:contato@samambai.com) when the sales page still does not state price and timeline, when the right robot is blocked, or when a competitor is the one getting cited for work you actually do. The 30-minute call is for naming which gate is shut. If the shut gate is "there is no llms.txt," the honest answer is that this is not the gate.
