- geo
- generative-ai
What gets a business cited by ChatGPT, Gemini and Google AI in 2026
In short
A business gets cited when the page is in that engine's index, answers the question with a price and a date, is corroborated by someone else, and stays current. An llms.txt file, special schema, and fabricated mentions do not clear those gates.
A business is cited by ChatGPT, Gemini, or Google's AI when it clears four gates, in this order: the page is in that engine's index, the page answers the question (with a price and a date when the question is commercial), someone outside the company confirms the business is real, and the page stays true. Google treats the whole job as SEO. In the guide updated on 10 July 2026, the line is plain: optimizing for generative search is optimizing for search Source [1].
The four gates are how samambai reads the official docs and the tests that hold up. They are not a Google product and they are not a GEO formula. A fifth gate (a secret file, special schema, a page chopped into chunks, a bought mention) is what Google told site owners to ignore Source [1].
The four gates, one sentence each
- Access to the index. If the right crawler cannot read the page, the page does not exist for that engine.
- A page that answers. Once pages are already retrieved, topic, an explicit price, and a recent date decide who is cited first. Formatting barely decides anything.
- Trust from other people. The answer reflects what the rest of the web says. A fabricated mention is the shortcut spam policy now names.
- Freshness. A stale fact loses to a current fact on a page that was already retrieved. A pile of new pages, one per query variant, is a different problem: spam.
Fail gate 1 and the rest is zero. Clear gate 1 and miss gate 2, and the engine knows the domain and cites a competitor. Gates 3 and 4 do not rescue a blocked site.
Gate 1: each engine has its own door
AI Overviews and AI Mode come out of the Google Search index. The guide describes two techniques. Retrieval-augmented generation (Google also calls it grounding): Search ranking retrieves pages, the model reads passages, and the answer shows clickable links. Query fan-out: the model fires several related searches at once. The official example is "how to fix a lawn that's full of weeds," which can become "best herbicides for lawns," "remove weeds without chemicals," and "how to prevent weeds in lawn" Source [1]. A single page titled "AI agency" does not cover those child searches by itself. Publishing one URL per variation does not fix it either. The same guide sends that tactic to the scaled content abuse policy Source [1].
For a supporting link to appear, the page has to be indexed and eligible for a snippet, and the site has to be included in the Search Console control for generative AI features. The default is include. Excluding the site removes its links and stops its content from grounding AI Overviews, AI Mode, and generative AI features in Discover. It does not change ranking in the rest of Search, and it does not control model training. The control reached every site on 31 August 2026. Exclusion usually applies within 1 to 2 days, with some cache lag Source [1] Source [10].
The Google-Extended token is a different switch, and mixing the two up is an expensive mistake. It has no user-agent of its own. Crawling still happens with the user-agents Google already uses. The token only decides whether content Google has already crawled may train future Gemini models and ground answers in the Gemini app and in Grounding with Google Search on Vertex AI. It does not change inclusion or ranking in Google Search Source [9]. Blocking Google-Extended does not remove a business from AI Mode. Allowing it does not put a business into AI Mode. The AI Mode door is Googlebot, plus indexing, plus the Search Console control.
| Engine | Where the page comes from | What to allow | What is not the door |
|---|---|---|---|
| AI Overviews and AI Mode | Google's index | Googlebot, snippet allowed, Search Console AI control set to include | Google-Extended, llms.txt |
| Gemini app and Vertex grounding | Already crawled content, under the token | A separate choice: allow or block Google-Extended | Treating this as the AI Mode switch |
| ChatGPT search | OAI-SearchBot, and sometimes partners | OAI-SearchBot and the IPs in openai.com/searchbot.json | GPTBot (that one is training) |
| Copilot and Bing AI summaries | Bing's index | Bing Webmaster Tools and, for freshness, IndexNow | Assuming Google is on IndexNow |
| Claude search | Claude-SearchBot, on the crawler page | Claude-SearchBot in robots.txt | Treating ClaudeBot (training) as search |
| Brave, and Anthropic's government connector | Brave's index; in government, the Brave API | A page Googlebot is allowed to crawl | A distinct Brave user-agent (Brave does not advertise one) |
On ChatGPT, the bots page separates four agents. Each setting is independent Source [4].
- OAI-SearchBot surfaces the site in ChatGPT search. It is not training. A site that opts out is not shown in search answers, though it can still appear as a navigational link. OpenAI asks for
Allowand for the host to accept the published IP ranges. A robots.txt change can take about 24 hours. If both OAI-SearchBot and GPTBot are allowed, OpenAI may use one crawl for both uses. - GPTBot may feed foundation-model training. Blocking GPTBot is a training decision. It is not the search switch.
- ChatGPT-User visits a page because a person asked. OpenAI says robots.txt rules may not apply, and that this agent is not used to decide whether content appears in Search.
- OAI-AdsBot visits only landing pages submitted as ads. It does not train foundation models.
The ChatGPT search help page, marked "updated last month" when read on 2 October 2026, says ChatGPT ranks results using multiple factors meant to find relevant, reliable information, and that placement is not guaranteed. The eligibility condition on the page is: allow OAI-SearchBot, and make sure the host or CDN allows traffic from the published IPs. The same page says search sometimes partners with other providers, rewrites the question, and points at Microsoft's and Shopify's privacy policies. That documents a partnership. It does not document that the index is Bing, and it does not document an overlap percentage Source [5].
Anthropic publishes three robots. ClaudeBot is training. Claude-User fetches pages when a person asks. Claude-SearchBot crawls to improve search results. Blocking the search bot can reduce visibility and accuracy in search answers. Opt-out is by robots.txt. Blocking an IP can stop the crawler from reading robots.txt, and Anthropic warns that an IP opt-out may not stick. The bots honor robots.txt and do not try to bypass CAPTCHAs Source [6].
Anthropic's crawler page does not name the index behind commercial Claude. A separate page, the web search connector for Claude for Government, says native web search is off in that product and the connector calls the Brave Search API. Only the query string goes to Brave. History, identity, and files do not Source [8]. "All Claude search is Brave" is not a sentence Anthropic wrote. The operational fact Brave does publish still matters: Brave's crawler does not advertise its own user-agent, and if Googlebot cannot crawl a page, Brave's bot will not crawl it either Source [7]. Shut Google's door and Brave's door shuts with it.
IndexNow, in the JSON read on 2 October 2026, notifies Bing, Yandex, Seznam, Naver, Yep, the Internet Archive, and Amazonbot. Google is not on the list Source [14]. A ping is not indexing. It is how Bing learns that a page changed, which is the use the AI Performance announcement recommends for freshness Source [12].
Gate 2: the page answers, with a price and a date
Google's guide asks for content that is not a commodity: a point of view you cannot recycle from another site or get from a generic model, first-hand experience, text organized for a person, and images or video when they actually help Source [1]. "Seven tips for hiring an AI vendor" is a commodity. "What it costs to take a Lovable app into production, and what breaks in a multi-tenant Supabase backend" is a page only the team that did the work can write.
Among pages already in front of the model, the cleanest 2026 test is by Vishwakarma, Kumar, and Jamidar, researchers at Sprinklr, submitted to arXiv on 25 May 2026 and presented at SIGIR 2026 Source [15]. It is not a tool vendor's blog post. It is an experiment. It is also not the open web, and that limit matters more than the large number.
They injected exactly two documents into the context, across 252,000 trials and six models (Gemini 2.5 Flash, GPT-5 Nano, GPT-5 Mini, GPT-5.2, Claude 3.5 Sonnet, and Kimi K2 Thinking). The two texts differed in one factor only. Brands were anonymized. Document order was swapped so content would not be confused with position bias. They never called a search engine. Retrieval was off.
Four factors were unanimous across all six models, with a huge effect (odds ratio above 100, which in this design means a near-deterministic preference in a two-page contest, not a 100x lift on Google):
- the text is about something else (topic mismatch)
- the price is missing
- the date is 2019 against a 2026 date
- the document is second in the injected list
Completeness and trust cues help less. Formatting-only edits (a dense paragraph versus sections, information scattered through the page) had no consistent effect. "Recent date versus no date" was not a consensus result either. What was consensus is a recent date against an old date Source [15].
The practical reading, with the lab limit in front: price and date do not get a page into the index. They break a tie between pages that were already retrieved. A service page with no price ("contact us") loses, in this test, to the page that states the number. A page dated 2019 loses to a page whose facts are from 2026. Stamping today's date on stale copy is not the factor they measured. They changed the dated content.
For a services firm, the page that clears this gate answers out loud, in the first paragraph:
- what is delivered (a WhatsApp agent that books appointments, an n8n automation, an app that outgrew Lovable)
- for whom (a clinic group, a hotel group, a B2B wholesaler)
- how long (most automations live in 2 to 4 weeks, a full project in 4 to 8 weeks)
- how much (a sprint at US$1,200 for two weeks, a project from US$3,000, a partnership from US$900 a month)
- the date those figures were written
Bing, in the 10 February 2026 AI Performance launch, recommends the same kind of page as product guidance, without a controlled test: depth, headings, tables, an FAQ in the body, evidence, freshness, and text that matches the images and video. Citations in the report do not indicate placement inside the answer Source [12]. Structure helps a person and helps extraction. It does not replace price, date, and topic.
Gate 3: trust comes from outside, and a fake mention is spam
The guide says generative answers can show what the web says about a product, in blogs, videos, and forums. It then takes apart the shortcut: chasing inauthentic mentions is less useful than it looks. Ranking systems look at quality. Spam systems block spam. Generative features depend on both Source [1].
Since 15 May 2026 that is policy, not advice. The changelog entry that day clarifies that spam policies also apply to generative AI responses in Google Search. It was not a new policy. It was the notice that the existing policy already covered AI Overviews and AI Mode Source [2]. The current policy text, updated on 28 August 2026, defines spam as a technique used to deceive users or to manipulate Search systems, "such as attempting to manipulate generative AI responses in Google Search" Source [3].
What that text covers, in buyer language:
- Scaled content. Many pages whose main purpose is to manipulate rankings, not to help. That includes generating pages with AI that add nothing, and standing up multiple sites to hide the scale Source [3]. One URL per city, or one URL per fan-out query, with the same text and a swapped city name, is the pattern the July guide points at that policy Source [1].
- Doorways. Pages built to rank for a query and then push the person somewhere else. Several domains with near-identical homepages. City pages that dump everyone onto one URL Source [3].
- Link spam. Buying links, requiring a link in a contract without letting the other party use
nofolloworsponsored, widgets and footers spread across many sites, low-value content built to manufacture a signal Source [3]. - Expired domains. Buying an old domain to inherit rankings with content that has little or no value Source [3].
Outside Google, Microsoft documented the commercial cousin of that shortcut. On 10 February 2026 the security team described "AI Recommendation Poisoning": "Summarize with AI" buttons that open an assistant with an instruction to remember the company as a trusted source or to recommend it first. Over 60 days of AI URLs in email traffic, they found more than 50 distinct prompts, from 31 companies, across 14 industries (finance, health, legal, SaaS, agencies, food, business services). Effectiveness varied by assistant and over time. Microsoft classifies the case as memory poisoning Source [19]. That is an attack. It is not a brand tactic.
Trust that clears gate 3 is dull and checkable. A partner-directory listing the company actually earned. A reported article. A customer page the customer controls. A person's profile, with a name and a history, that matches the site. For a local business, Google points to Business Profile and Merchant Center as the path into AI answers Source [1]. Bing points to Bing Places for address, hours, and contact on local answers Source [12]. A software studio that has done work in São Paulo, New York, Madrid, and Tel Aviv does not become "the best agency in town" with a fake local listing in twenty cities. It becomes a true listing where it actually operates, plus pages that survive a check.
Gate 4: freshness is a new fact, not a publishing calendar
In the SIGIR test, a 2026 date against a 2019 date was one of the four unanimous factors, with retrieval turned off Source [15]. On Bing, the AI Performance announcement treats updates as the condition for an answer to use the current version, and points to IndexNow as the ping to participating engines Source [12]. Google states the other side: a high quantity of pages does not make a site more relevant, and the systems already understand relevance without an exact word match Source [1].
The freshness work that fits a services firm is short.
- When the price changes, the pricing page changes the same day. A sprint still listed at an old number teaches the model the wrong number. The public figure is US$1,200 per two-week sprint.
- When the timeline changes, the sentence "2 to 4 weeks" changes. A new post is the wrong tool for fixing one sentence.
- The visible date follows a revision of the facts, not a cosmetic title tweak.
- One good page, kept current, beats twelve thin pages covering the fan-out.
What does not move citations
The 10 July 2026 guide has a mythbusting section. Four items match what agencies still sell Source [1]. The 15 June 2026 changelog had already added the llms.txt note: the file is not needed for Google Search, and it does not change visibility or rankings in either direction Source [2].
| Tactic | What the source says | Note |
|---|---|---|
| llms.txt | Google Search does not use it. Keeping the file neither helps nor hurts on Google Source [1]. | Vendor study: 97% of valid files received zero requests in May 2026 Source [16]. Detail in the llms.txt piece. |
| Special AI schema | Not required for generative search. There is no special schema.org markup for it. Schema remains useful for classic rich results Source [1]. | Vendor, difference-in-differences: 1,885 pages that gained JSON-LD between August 2025 and March 2026, against about 4,000 controls. AI Mode up 2.4% and ChatGPT up 2.2%, indistinguishable from zero. AI Overviews down 4.6%, a small effect. The pages were already heavily cited. The test does not speak for pages that are not cited yet Source [17]. |
| Chunking | There is no requirement to break the page into tiny pieces. Google says it can understand several topics on one page Source [1]. | In the SIGIR lab, formatting alone had no consistent effect Source [15]. |
| Inauthentic mentions | The guide says to ignore them. Spam policy, since 15 May 2026, covers manipulating an AI answer Source [1] Source [2] Source [3]. | A "remember us" button was documented by Microsoft as memory poisoning Source [19]. |
| Rewriting only for the AI | Not required. The systems understand synonyms and intent Source [1]. | A page written for a person, with a checkable fact, is what the guide asks for. |
Google's FAQ rich result was deprecated in the 8 May 2026 changelog entry Source [2]. An FAQ in the body of the page is still text. An FAQ badge in classic search is not a citation lever.
What has counted as spam since 15 May 2026
That date did not invent a new penalty with GEO in the name. It made explicit that manipulating a generative AI response is Search spam Source [2] Source [3]. In practice, these jobs now have a name in the policy if the main purpose is manipulation:
- a network of satellite sites citing each other
- dozens of near-duplicate pages, one per fan-out variant or per city
- hidden text with an instruction to the model
- a link required by contract, without
rel="sponsored"orrel="nofollow" - an expired domain stuffed with commercial pages that add nothing
Detection is automated and, when needed, human review that can end in a manual action. The site can rank worse or drop out of results Source [3]. "An agency published it" is not a defense written into the policy.
How to measure this without fooling yourself
Position inside one AI answer is not a metric. Presence, across many runs, is what the official instruments and the consistency test can support.
Search Console has a generative AI performance report with impressions of links in AI Overviews and AI Mode: pages, countries, dates (Pacific Time), and devices. Since 31 August 2026 the report is available to sites worldwide, if there are enough impressions. The help page, as opened, has no columns for clicks, CTR, position, or query. The same data also sits inside the Web performance report. Adding one on top of the other counts the same impression twice Source [11].
Bing Webmaster Tools has shown citations since 10 February 2026: total citations, average cited pages per day, a sample of grounding queries, citations by URL, and the time series. The announcement says, in more than one place, that this does not indicate placement, ranking, or the role of the page inside a single answer Source [12]. On 16 June 2026, Intents, Topics, Citation Share, and Compare entered a global preview. Citation Share is the percentage of citations for that grounding query that went to your site. Microsoft writes that it is not a ranking, not a traffic share, does not reveal a competitor's domain, and does not assign a quality score Source [13].
The consistency test is Rand Fishkin's, at SparkToro, with collection in November and December 2025. About 600 volunteers ran 12 prompts on ChatGPT, Claude, and Google's AI, 2,961 times in total. The chance that the same brand list appears in two answers was under 1 in 100. The same list in the same order was about 1 in 1,000. What repeated was the set of brands that show up often (a presence rate), not the order. The data partner works at Gumshoe, an AI tracking company. The write-up treats position as a metric with no foundation, and presence across many runs as the measure that remains Source [18].
What is not a measurement:
- a screenshot of one chat
- "we came in second" in a single answer
- a third-party tool that promises an internal Google metric (the guide says no third party has access to those systems) Source [1]
- adding generative AI impressions on top of Web performance Source [11]
An honest cadence, once Search Console and Bing are verified: look at both reports every week, and run a fixed set of real customer questions many times a month, recording whether the brand appeared. The trend of that set is the result. One day's swing is not.
Common mistakes
- Blocking GPTBot and concluding the site "left ChatGPT." The search switch is OAI-SearchBot Source [4].
- Allowing the bot in robots.txt and blocking the IP at the edge. OpenAI conditions eligibility on both Source [5].
- Treating
Google-Extendedas the AI Mode switch Source [9]. - Publishing twenty cities with no operation, to "cover the fan-out" Source [1] Source [3].
- Hiding the price and hoping the model invents a flattering number. In the lab, a missing price was disqualifying Source [15].
- Measuring a screenshot and calling it a rank Source [18].
- Buying an llms.txt generator and calling that the program Source [1] Source [16].
What not to do
Do not buy mentions. Do not build a domain network. Do not hide an instruction to the model. Do not publish a page per synonym. Do not rewrite the site "in the AI's tone." Do not swap useful copy for FAQ markup aimed at a rich result. Do not promise, and do not accept a promise of, first place in ChatGPT.
If the true page does not exist yet, the work is to write one: the service, the price, the timeline, the date, an anonymized example of work already shipped. A multi-tenant B2B SaaS for gyms that started in Lovable and now runs in production with paying customers. WhatsApp agents in production for a clinic group, for hospitality and B2B retail, for an adventure park, for a property manager, and for a visa advisory. That is gate 2 material. Ten variations of the same paragraph are not.
When not to hire samambai
Do not hire samambai if the ask is any of these:
- a guarantee of position, a citation percentage, or "first in ChatGPT in 30 days"
- a site network, a paid mention, or a button that asks the model to memorize the brand
- an llms.txt file, a schema install, or "chunking" as the whole deliverable
- paid media management (not part of what samambai sells)
- a result next week from a domain Google and Bing have not indexed, with no new page and no willingness to wait for a crawl
The work fits when there is a real service, a page that can state the price, and patience to measure presence. The first conversation is 30 minutes, with no commitment. The public formats are on pricing: a sprint at US$1,200 for two weeks, cancel any time; a project from US$3,000, with scope, timeline, and price fixed, and a delivery every sprint; a partnership from US$900 a month. The GEO service is described on the GEO page. Why the llms.txt file is not this work is in the companion piece.
The criterion is the same one used on every engagement: diagnosis, strategy, execution, and evolution, in two-week sprints, with a human review. AI agents sit inside the build, the tests, and a separate audit. They do not replace someone reading the page before it goes live.
Questions
Does ranking first on Google get you cited in ChatGPT?
No. Google uses its own index for AI Overviews and AI Mode, and indexing still does not guarantee display. ChatGPT uses OAI-SearchBot and, sometimes, search partners. OpenAI's help page says placement is not guaranteed.
Will publishing llms.txt get the company into ChatGPT?
Not on the evidence Google and OpenAI publish. Google says it ignores the file. OpenAI's crawler docs tell you to allow OAI-SearchBot and the published IP ranges. They do not mention llms.txt.
Can a vendor guarantee the first citation?
No. In a volunteer test, the same brand list showed up in fewer than 1 in 100 answers. A promise of first place is a screenshot, not a measurement.
Do schema, FAQ markup, or chunking raise citations?
Google says generative search needs no special schema, and the guide says to ignore chunking. In a vendor study of 1,885 pages, adding JSON-LD did not move ChatGPT or AI Mode citations by an amount you can tell from zero.
How do you measure this without fooling yourself with a screenshot?
Track presence across many runs of the same question, plus impressions in Search Console's generative AI report and citations in Bing AI Performance. One screenshot is an anecdote.
Sources
- Optimizing your website for generative AI features on Google Search. Google Search Central. . back to text
- Latest Google Search documentation updates. Google Search Central. . back to text
- Spam policies for Google web search. Google Search Central. . back to text
- Overview of OpenAI crawlers. OpenAI. . back to text
- Searching the web with ChatGPT. OpenAI Help Center. . back to text
- Does Anthropic crawl data from the web?. Anthropic. . back to text
- Brave Search Crawler. Brave. . back to text
- MCP: Web Search (Claude for Government). Anthropic. . back to text
- List of Google's common crawlers. Google. . back to text
- Search generative AI control. Google Search Console Help. . back to text
- Generative AI performance report (Search). Google Search Console Help. . back to text
- Introducing AI Performance in Bing Webmaster Tools (Public Preview). Microsoft Bing. . back to text
- New AI Visibility Insights in Bing Webmaster Tools. Microsoft Bing. . back to text
- IndexNow participating search engines. IndexNow. . back to text
- What Gets Cited: Competitive GEO in AI Answer Engines. arXiv, SIGIR 2026. . back to text
- We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read. Ahrefs. . back to text
- We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.. Ahrefs. . back to text
- AIs are highly inconsistent when recommending brands or products. SparkToro. . back to text
- Manipulating AI memory for profit: The rise of AI Recommendation Poisoning. Microsoft Security. . back to text
Find out whether AI already recommends your company.
A 30-minute call, no strings attached.
or write to contato@samambai.com