• geo
  • llms-txt

Does llms.txt get you cited by ChatGPT? What Google and OpenAI actually say

In short

llms.txt does not get a business cited by ChatGPT, Gemini, or Google's AI. Google says Search ignores the file. OpenAI tells sites to allow OAI-SearchBot and the published IP ranges. The file still works as an index for coding agents and docs.

No. Publishing an llms.txt file does not get a business cited by ChatGPT, Gemini, or Google's AI. Google says Search does not use the file, and that keeping one neither helps nor hurts visibility or rankings Source [1] Source [2]. On the page where OpenAI explains how a site shows up in ChatGPT search, the documented step is different: allow OAI-SearchBot and allow the published IP ranges. That page does not list llms.txt as a condition for a citation Source [4]. The file is still useful in a narrow case, a coding agent that needs a map of the docs. That is why samambai keeps one.

A vendor that sells the file as a citation shortcut is selling an index. An index is not an answer, not a third-party source, and not a rank.

What the file actually is

llms.txt is a proposal, not a required search standard. Jeremy Howard published the first version on 3 September 2024 and updated it to v2 on 10 August 2026 Source [3]. The point is to hand an agent a short Markdown map instead of making it parse a full HTML page, with the nav, the ads, and the scripts.

The spec is small. The only required section is an H1 with the name of the project or the site. After that you can add a blockquote summary, free text that does not open new headings, and H2 sections whose bodies are lists of links. Each list item is a link, then an optional note after a colon. A section titled Optional is the convention for links an agent may skip when the context window is tight Source [3].

The file can sit at the root (/llms.txt) or on any path, and it covers the URLs under that path. If two files apply, the spec says to use the more specific one. It sits next to robots.txt and a sitemap. It does not replace them. robots.txt says what a robot may fetch. A sitemap lists pages for indexing. llms.txt is read on demand, when an agent needs a topic while helping a person. Howard's expectation is inference, not training, though a training run could use the text too Source [3].

The proposal is plain about where the file is used most: software documentation, where a coding agent follows links to an API reference and a tutorial. The same shape can, in his words, describe a business, a personal site, or a school. That is a possible use. It is not a promise that ChatGPT will cite the business Source [3].

The proposal also asks for a clean Markdown twin of important pages, at the same URL with .md appended, plus rel="alternate" links or an HTTP Link header. Chrome Lighthouse, the spec says, audits for the file as part of its agentic browsing checks Source [3]. A developer-tool audit is not a Google ranking factor. Google's own guide, in the mythbusting section, tells site owners to ignore AI text files Source [1].

What Google has said

"Optimizing your website for generative AI features on Google Search" was last updated on 10 July 2026. The mythbusting section says you do not need new machine-readable files, AI text files, markup, or Markdown to appear in Google Search, including its generative AI features, because Google Search itself does not use them. Discovering, crawling, and indexing a file does not mean the file gets special treatment Source [1].

The next sentences are the ones that matter if you are afraid to delete the file. It is fine to create and maintain an llms.txt (or a similar file) for other services or systems that use it. Doing so will neither harm nor help your visibility or rankings in Google Search, because Google Search ignores those files Source [1].

That note was added on 15 June 2026. The Search Central changelog records the reason in the same terms: the files are not needed for Google Search, and they will not negatively or positively affect visibility or rankings. Keeping them for other services or systems is fine Source [2]. This is not a paraphrase from a conference hallway. It is the official documentation changelog.

The same guide tells you to ignore, in the same breath, chopping pages into tiny pieces ("chunking"), rewriting copy only for AI systems, chasing inauthentic mentions, and treating a special schema type as a requirement for generative search. Structured data remains useful for classic rich results. There is no special schema.org markup for AI Overviews or AI Mode Source [1]. The full citation model, index then page then third-party trust then freshness, is in what gets a business cited by AI.

On 15 May 2026 the same changelog clarified that existing spam policies also apply to generative AI responses in Google Search Source [2]. Stuffing hidden instructions into llms.txt, or publishing dozens of pages only to cover query variants, is not "file optimization." It sits under the spam rules Google says already cover the generated answer. The guide warns that separate pages for every variation of a query, built to manipulate rankings or generative responses, violate the scaled content abuse policy Source [1].

What a log study actually measured

An official statement answers "does Google use this?" A log answers "did anyone request the file?" Those are different questions. The second one has a vendor study, not a Google measurement.

On 15 June 2026, Louise Linehan and Xibeijia Guan at Ahrefs published a read of Ahrefs Web Analytics and Bot Analytics Source [6]. The population was every domain in that product with traffic in May 2026: 137,210 domains. Ahrefs says those customers skew more technical and more SEO-aware than the web at large, so the adoption rate is an upper bound, not an internet average.

They looked for /llms.txt at the root returning HTTP 200, checked that the body was Markdown rather than an HTML error page, and counted requests to that path. They did not check whether each file matched the spec. The study measures the index file, and only the index file Source [6].

What they measuredFigure in the Ahrefs studyHow to read it
Domains with traffic in May 2026137,210Vendor base, skewed toward technical customers
Publish a valid llms.txt28% (38,360)A ceiling, not a web average
Files with zero requests in May 202697%No bot, no person
Requests that did arrive and came from bots96%The other 4% were people
Requests from named AI tools, among the ones that arrived19.5%GPTBot first, Claude-Code second
AI search retrieval bots (OAI-SearchBot, PerplexityBot, Claude search)1.1% of requestsThe bots that cite barely read the index
Requests from people studying the file (GEO tools, scanners, research)about 12%Audit traffic, not citations
AI-bot requests to an /llms.txt that 404szeroNothing goes looking for a file you did not publish

Ahrefs' chart marks 38,360 domains with a valid file, 28% of the 137,210 base. The complement drawn on that same chart, 98,640, adds up to 137,000 with the 38,360. That is the chart rounding, not a second count. The roughly 3% that received any request hold the measured traffic, on the order of 1,100 domains and about 22,000 requests. Inside that small pool, Chrome's Lighthouse llms.txt audit accounted for about 1 in 1,000 fetches (22 requests) Source [6].

Two details should change a buying decision. First, "fetched" is not "used." Ahrefs writes that many bots may have downloaded the file without acting on it, so every AI percentage is a ceiling Source [6]. Second, Slack's link-preview bot fetched the file more often than PerplexityBot. A chat unfurl is not a citation.

Ahrefs also splits the readers. Agents and agent infrastructure were 10.5% of requests. Training crawlers were 5.3%, with GPTBot alone at 4.51% and ClaudeBot at 0.8%. User-triggered assistants were 2.5%. Retrieval bots were 1.1%, and inside that group OAI-SearchBot led at 0.74%. Claude-Code, Anthropic's coding agent, out-fetched every retrieval bot, every assistant, and every training crawler except GPTBot and an indexer named statespace-indexer, whose operator Ahrefs names without confirming the IP ranges Source [6].

Googlebot showed up about 900 times in May. Ahrefs reads that the way the guide does: Googlebot fetches URLs it discovers, the same way it fetches a sitemap. Those fetches do not show special interest in llms.txt, and the study cannot see whether any of that text later feeds Gemini Source [6]. That matches the guide. Indexing one more file type does not create special treatment Source [1].

On paths that returned 404, the split flipped. A valid file drew 96% bot traffic. A missing file drew 98% human traffic, and the AI-bot share of those 404s was zero. The person typing the URL into a browser, usually to check a competitor, is a person. An AI system does not hunt for an llms.txt you never published Source [6].

One security finding is worth keeping, without the drama. Among research bots (2.7% of requests), the largest identified itself as prompt-injection-survey/1.0. Someone is scanning the file as a prompt-injection surface, because agents are built to trust what the index says Source [6]. Reason enough to treat the file as public, short text. Not a reason to hide instructions in it.

A vendor study measures that vendor's logs. It does not measure how often ChatGPT cited a brand. It does not prove that deleting the file drops citations, or that publishing one raises them. It proves something simpler. In May 2026, almost nobody requested the file, and the requests that existed were rarely from the search bot that cites.

What OpenAI actually documents

"Overview of OpenAI crawlers," as read on 2 October 2026, separates four agents. Each control is independent Source [4].

AgentWhat OpenAI says it is forWhat it is not
OAI-SearchBotSurfaces sites in ChatGPT search results. A site that opts out is not shown in search answers, though it can still appear as a navigational link.Not the llms.txt file
GPTBotCrawls content that may be used to train foundation models. Disallowing it says the content should not train those models.Does not decide search
ChatGPT-UserA visit triggered by a person in ChatGPT or a custom GPT. Because the person started the action, robots.txt rules may not apply.OpenAI says it does not decide Search inclusion
OAI-AdsBotVisits only pages submitted as ads, to check the page and the ad's relevance. The data is not used to train foundation models.Not organic inclusion

The written recommendation for search is specific: allow OAI-SearchBot in robots.txt, and allow requests from the IP ranges published at https://openai.com/searchbot.json Source [4]. GPTBot, ChatGPT-User, and OAI-AdsBot have their own lists. If a site allows both OAI-SearchBot and GPTBot, OpenAI may use a single crawl for both jobs so it does not crawl twice. A robots.txt change takes about 24 hours to propagate in their systems Source [4].

The sample OAI-SearchBot user agent on that page ends in OAI-SearchBot/1.4 and points at https://openai.com/searchbot. When they fetch robots.txt, they may add a robots.txt marker to the user agent so logs can tell that fetch from the others Source [4]. The version number can change. The page says so.

Nothing on that page says "publish an llms.txt." The only mention of the filename is in the docs header, and the meaning is different: "For the complete documentation index, see llms.txt" Source [4]. That is an index of developer documentation. The file at https://developers.openai.com/llms.txt, fetched the same day, confirms it. The H1 is "OpenAI Developers." The blockquote describes the hub for the API, Ads, Codex, and the rest. The H2 sections point at other product llms.txt files and at .md guides: quickstart, reference, evals, agentic commerce Source [5]. That is the spec's format, used so an agent can find the right docs page. It is not a file telling ChatGPT to cite OpenAI when someone asks for a vendor.

Mixing up the four robots is the expensive mistake. Blocking GPTBot is not the same as leaving ChatGPT search. Allowing GPTBot does not put the page in the answer. ChatGPT-User may visit a URL a person asked for, and that visit is not llms.txt being "read by ChatGPT." Ads have their own robot and, in the docs, do not train foundation models. Showing up in a search answer follows the OAI-SearchBot line and the IP allowlist, and even then a citation is not guaranteed. The page still has to be retrieved, and the model still has to use it. The index file is not in that sentence.

When the file is still worth having

It is worth having when someone will point an agent at the site and that agent needs a map. The spec describes the case: library docs, an API reference, a tutorial Source [3]. OpenAI does exactly that for developer documentation Source [4] Source [5]. The proposal says Anthropic's and Gemini's docs do the same, and it prints those URLs Source [3]. We did not fetch those two files for this article. What we did fetch is OpenAI's file and the spec.

It is also worth having as an internal index. A coding agent maintaining the site, or an agent inside the operation, spends less context if it starts from a short list of Markdown pages than if it starts from the HTML. The spec says to test the file by asking an agent questions with only that llms.txt as the starting point Source [3]. That test tells you whether the map is clear. It does not tell you whether you will be cited.

It is not worth having as a visibility project. The vendor base rate is blunt: 97% of existing files had no reader of any kind for a full month, and the bots that retrieve answers were 1.1% of the requests that existed Source [6]. Publishing the file does not put the domain on a radar. The 404s show the opposite Source [6].

Why samambai keeps one

The studio publishes https://samambai.com/llms.txt. The file is a short map: what the studio does, public prices in US dollars, how an engagement runs, how to get in touch, and links to Markdown versions of the key pages. The reason is the same one behind OpenAI's developer docs. An agent that has already been sent to the site can read the service, the timeline, and the price without guessing at the HTML. The studio's own operation is AI-first: people plus agents for building, testing, and an independent audit, always with a human review. That internal agent is a plausible reader. So is a buyer who pastes the URL into a coding agent.

The file is not there to win a citation. Google says keeping it for other systems does not change Search Source [1] Source [2]. The vendor study says that when a reader shows up, it looks more like a coding agent than like a search bot Source [6]. Treating our file as proof of GEO would repeat the myth this page takes apart.

If the map goes stale, it lies to the only reader that matters. Price, timeline, and scope in llms.txt have to match the site. The spec asks for short language and a note on each link Source [3]. A stale index is worse than no index, because the agent that trusts it repeats the error. If the site says US$1,200 for a sprint and the file still says US$800, the agent quotes US$800.

What to do instead

Citation follows a different path. The long version is what gets a business cited by ChatGPT, Gemini, and Google's AI. The operational order, the order in which the work actually stalls, is short.

  1. Access to the index of the engine you care about. On Google, the page has to be indexable and eligible for a snippet, and the site has to be included in Search Console's generative AI control. The guide says meeting the requirements does not guarantee indexing or serving Source [1]. On ChatGPT, search asks you to allow OAI-SearchBot and to allow the IP ranges in searchbot.json at the firewall or the CDN Source [4]. Blocking Googlebot "so they cannot train on us" and then expecting a citation in Google's AI contradicts the guide: Google's generative search is grounded in the Search index Source [1].
  2. A page that answers the buying question. Price, timeline, what is included, and the date that fact is true. The guide wants non-commodity content: something you know because you did the work, not a summary of what every other page already says Source [1]. "AI solutions for your business" does not answer "what does a two-week automation sprint cost."
  3. Trust that does not come from your own site. Bought mentions, a network of satellite domains, and hidden text in a "summarize with AI" button sit on the wrong side of the policy Google says covers the generated answer, clarified on 15 May 2026 Source [2]. The guide says chasing inauthentic mentions does not help, because core ranking and spam systems still apply to the generative feature Source [1].
  4. Real freshness, not a fresh stamp. If the price changed, the page changes. Putting today's date on last year's copy is not an update.

Measure without fooling yourself. Track presence across many runs of the same question, and impressions in Search Console's generative AI report. The guide points at that report. It shows discovery, not a rank Source [1]. One screenshot is an anecdote. A third-party tool that promises internal Google metrics does not have them. The guide says so outright Source [1].

GEO at samambai is that work: find out what is already in the index, make the page answer, earn real corroboration, and keep a rhythm of updates. It is not an llms.txt install. Public prices are on pricing: a Sprint at US$1,200 per two weeks, cancel whenever you want; a Project from US$3,000 with scope, timeline, and price fixed up front; a Partnership from US$900 a month. The first call is 30 minutes and creates no commitment. GEO for a specific site is quoted after that call, because it depends on what is already indexed. The service page is GEO.

A buyer in New York asking which integrator to hire, a buyer in Madrid comparing WhatsApp agents, and a buyer in Tel Aviv checking a fixed project price all hit the same wall. If the page does not state the price and the date, the index has nothing solid to repeat. The file at the root does not fill that gap.

Mistakes that waste a week

MistakeWhy it fails
Publish the file and expect a citation within a week97% of files in the study got no request in a month. Search bots were 1.1% of what did arrive Source [6].
Assume Google rewards or punishes the fileThe guide and the 15 June 2026 changelog say the opposite: ignored, no help, no harm Source [1] Source [2].
Block GPTBot and conclude you have left ChatGPTThe controls are independent. Search is OAI-SearchBot. Training is GPTBot Source [4].
Allow only ChatGPT-UserOpenAI says that agent does not decide Search inclusion Source [4].
Generate the file from a plugin and never touch it againThe plausible reader is an agent that trusts the map. A stale map repeats the wrong price.
Count a fetch in the log as a winAhrefs says a fetch is not use Source [6].
One URL per query variant, then list them all in the fileThe guide treats volume built to manipulate the answer as scaled content abuse Source [1].
Hide an instruction that tells the model to store a memoryThe file is public. The study already saw a crawler identifying itself as prompt-injection research Source [6].

What not to do

Do not swap the page that answers for an index file. Do not buy mentions. Do not stand up a domain network whose only job is to "cite" the brand. Do not promise, and do not hire anyone who promises, the first slot in ChatGPT. Do not use llms.txt as a place to instruct the model. Do not treat schema, chunking, or a cosmetic rewrite as a substitute for the index and for proof Source [1].

Do not delete the file because you are afraid Google will punish it. There is no documented penalty, and no documented prize Source [1] Source [2]. Delete it, or skip creating it, when nobody will point an agent at it and the cost of keeping it true is real. Fifteen lines that state the wrong price cost more than zero lines.

Do not read the Ahrefs study as if it were Google. It is one vendor's logs, on a technical customer base, for one month (May 2026). It is enough to retire the idea that "everyone is already reading llms.txt." It is not a citation rate.

When not to hire samambai

Do not hire the studio if the ask is one of these:

  • A guarantee of first place in ChatGPT, Gemini, or AI Mode. Nobody reading the docs can offer that. Google's guide says indexing does not guarantee serving Source [1].
  • Paid media. Ad management is not part of what samambai sells. OAI-AdsBot is OpenAI's product for people who advertise in ChatGPT, and it is not work this studio sells Source [4].
  • A site network, a mention package, or a "summarize with AI" button that tries to write memory into the assistant. That is the kind of manipulation Google's spam policy reaches inside the generated answer Source [2].
  • A project whose deliverable is the llms.txt file, a special schema type, or a page chopped into chunks. Google says to ignore all three Source [1].
  • A result in a few days on a site that is not in the index and has no third-party corroboration. Without the first gate, the file does not create the second.
  • A new URL for every way of asking the same question. The guide calls that a weak strategy and, when the goal is manipulation, scaled content abuse Source [1].

Write to contato@samambai.com when the sales page still does not state price and timeline, when the right robot is blocked, or when a competitor is the one getting cited for work you actually do. The 30-minute call is for naming which gate is shut. If the shut gate is "there is no llms.txt," the honest answer is that this is not the gate.

Questions

If I publish llms.txt, will ChatGPT start citing my company?

No. OpenAI's crawler docs say ChatGPT search depends on allowing OAI-SearchBot and the IP ranges at openai.com/searchbot.json. That page does not list llms.txt as a search requirement.

Will Google penalize my site for having the file?

No. The guide updated on 10 July 2026 says Google Search does not use the file, and that keeping it neither helps nor hurts visibility or rankings, including in generative AI features.

My server logs never show a hit on llms.txt. Is the file broken?

Usually nobody asked for it. In Ahrefs' vendor study, 97% of valid files got zero requests in May 2026. An empty log is not a writing problem.

Why does samambai still publish one?

It is a short index for an agent that has already been pointed at the site: services, public prices, and Markdown pages. It is not a request to be cited in ChatGPT.

What should I do instead if I want to be cited?

Clear four gates: get into the index, answer the question with a price and a date, earn third-party corroboration, and keep the page current. The full version is on the citation guide.

Sources

  1. Optimizing your website for generative AI features on Google Search. Google Search Central. . back to text
  2. Clarifying guidance on llms.txt files. Google Search Central. . back to text
  3. The /llms.txt file, v2. Jeremy Howard, llmstxt.org. . back to text
  4. Overview of OpenAI crawlers. OpenAI. . back to text
  5. OpenAI Developers llms.txt. OpenAI. . back to text
  6. We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read. Ahrefs. . back to text

Find out whether AI already recommends your company.

A 30-minute call, no strings attached.

or write to contato@samambai.com