---
title: "What gets a business cited by ChatGPT, Gemini and Google AI in 2026"
url: https://samambai.com/blog/what-gets-a-business-cited-by-ai/
date: 2026-10-03
author: samambai
---

A business gets cited when the page is in that engine's index, answers the question with a price and a date, is corroborated by someone else, and stays current. An llms.txt file, special schema, and fabricated mentions do not clear those gates.

A business is cited by ChatGPT, Gemini, or Google's AI when it clears four gates, in this order: the page is in that engine's index, the page answers the question (with a price and a date when the question is commercial), someone outside the company confirms the business is real, and the page stays true. Google treats the whole job as SEO. In the guide updated on 10 July 2026, the line is plain: optimizing for generative search is optimizing for search [1].

The four gates are how samambai reads the official docs and the tests that hold up. They are not a Google product and they are not a GEO formula. A fifth gate (a secret file, special schema, a page chopped into chunks, a bought mention) is what Google told site owners to ignore [1].

## The four gates, one sentence each

1. **Access to the index.** If the right crawler cannot read the page, the page does not exist for that engine.
2. **A page that answers.** Once pages are already retrieved, topic, an explicit price, and a recent date decide who is cited first. Formatting barely decides anything.
3. **Trust from other people.** The answer reflects what the rest of the web says. A fabricated mention is the shortcut spam policy now names.
4. **Freshness.** A stale fact loses to a current fact on a page that was already retrieved. A pile of new pages, one per query variant, is a different problem: spam.

Fail gate 1 and the rest is zero. Clear gate 1 and miss gate 2, and the engine knows the domain and cites a competitor. Gates 3 and 4 do not rescue a blocked site.

## Gate 1: each engine has its own door

AI Overviews and AI Mode come out of the Google Search index. The guide describes two techniques. Retrieval-augmented generation (Google also calls it grounding): Search ranking retrieves pages, the model reads passages, and the answer shows clickable links. Query fan-out: the model fires several related searches at once. The official example is "how to fix a lawn that's full of weeds," which can become "best herbicides for lawns," "remove weeds without chemicals," and "how to prevent weeds in lawn" [1]. A single page titled "AI agency" does not cover those child searches by itself. Publishing one URL per variation does not fix it either. The same guide sends that tactic to the scaled content abuse policy [1].

For a supporting link to appear, the page has to be indexed and eligible for a snippet, and the site has to be included in the Search Console control for generative AI features. The default is include. Excluding the site removes its links and stops its content from grounding AI Overviews, AI Mode, and generative AI features in Discover. It does not change ranking in the rest of Search, and it does not control model training. The control reached every site on 31 August 2026. Exclusion usually applies within 1 to 2 days, with some cache lag [1] [10].

The `Google-Extended` token is a different switch, and mixing the two up is an expensive mistake. It has no user-agent of its own. Crawling still happens with the user-agents Google already uses. The token only decides whether content Google has already crawled may train future Gemini models and ground answers in the Gemini app and in Grounding with Google Search on Vertex AI. It does not change inclusion or ranking in Google Search [9]. Blocking `Google-Extended` does not remove a business from AI Mode. Allowing it does not put a business into AI Mode. The AI Mode door is Googlebot, plus indexing, plus the Search Console control.

| Engine | Where the page comes from | What to allow | What is not the door |
|---|---|---|---|
| AI Overviews and AI Mode | Google's index | Googlebot, snippet allowed, Search Console AI control set to include | `Google-Extended`, llms.txt |
| Gemini app and Vertex grounding | Already crawled content, under the token | A separate choice: allow or block `Google-Extended` | Treating this as the AI Mode switch |
| ChatGPT search | OAI-SearchBot, and sometimes partners | OAI-SearchBot and the IPs in `openai.com/searchbot.json` | GPTBot (that one is training) |
| Copilot and Bing AI summaries | Bing's index | Bing Webmaster Tools and, for freshness, IndexNow | Assuming Google is on IndexNow |
| Claude search | Claude-SearchBot, on the crawler page | Claude-SearchBot in robots.txt | Treating ClaudeBot (training) as search |
| Brave, and Anthropic's government connector | Brave's index; in government, the Brave API | A page Googlebot is allowed to crawl | A distinct Brave user-agent (Brave does not advertise one) |

On ChatGPT, the bots page separates four agents. Each setting is independent [4].

- **OAI-SearchBot** surfaces the site in ChatGPT search. It is not training. A site that opts out is not shown in search answers, though it can still appear as a navigational link. OpenAI asks for `Allow` and for the host to accept the published IP ranges. A robots.txt change can take about 24 hours. If both OAI-SearchBot and GPTBot are allowed, OpenAI may use one crawl for both uses.
- **GPTBot** may feed foundation-model training. Blocking GPTBot is a training decision. It is not the search switch.
- **ChatGPT-User** visits a page because a person asked. OpenAI says robots.txt rules may not apply, and that this agent is not used to decide whether content appears in Search.
- **OAI-AdsBot** visits only landing pages submitted as ads. It does not train foundation models.

The ChatGPT search help page, marked "updated last month" when read on 2 October 2026, says ChatGPT ranks results using multiple factors meant to find relevant, reliable information, and that placement is not guaranteed. The eligibility condition on the page is: allow OAI-SearchBot, and make sure the host or CDN allows traffic from the published IPs. The same page says search sometimes partners with other providers, rewrites the question, and points at Microsoft's and Shopify's privacy policies. That documents a partnership. It does not document that the index is Bing, and it does not document an overlap percentage [5].

Anthropic publishes three robots. ClaudeBot is training. Claude-User fetches pages when a person asks. Claude-SearchBot crawls to improve search results. Blocking the search bot can reduce visibility and accuracy in search answers. Opt-out is by robots.txt. Blocking an IP can stop the crawler from reading robots.txt, and Anthropic warns that an IP opt-out may not stick. The bots honor robots.txt and do not try to bypass CAPTCHAs [6].

Anthropic's crawler page does not name the index behind commercial Claude. A separate page, the web search connector for Claude for Government, says native web search is off in that product and the connector calls the Brave Search API. Only the query string goes to Brave. History, identity, and files do not [8]. "All Claude search is Brave" is not a sentence Anthropic wrote. The operational fact Brave does publish still matters: Brave's crawler does not advertise its own user-agent, and if Googlebot cannot crawl a page, Brave's bot will not crawl it either [7]. Shut Google's door and Brave's door shuts with it.

IndexNow, in the JSON read on 2 October 2026, notifies Bing, Yandex, Seznam, Naver, Yep, the Internet Archive, and Amazonbot. Google is not on the list [14]. A ping is not indexing. It is how Bing learns that a page changed, which is the use the AI Performance announcement recommends for freshness [12].

## Gate 2: the page answers, with a price and a date

Google's guide asks for content that is not a commodity: a point of view you cannot recycle from another site or get from a generic model, first-hand experience, text organized for a person, and images or video when they actually help [1]. "Seven tips for hiring an AI vendor" is a commodity. "What it costs to take a Lovable app into production, and what breaks in a multi-tenant Supabase backend" is a page only the team that did the work can write.

Among pages already in front of the model, the cleanest 2026 test is by Vishwakarma, Kumar, and Jamidar, researchers at Sprinklr, submitted to arXiv on 25 May 2026 and presented at SIGIR 2026 [15]. It is not a tool vendor's blog post. It is an experiment. It is also not the open web, and that limit matters more than the large number.

They injected exactly two documents into the context, across 252,000 trials and six models (Gemini 2.5 Flash, GPT-5 Nano, GPT-5 Mini, GPT-5.2, Claude 3.5 Sonnet, and Kimi K2 Thinking). The two texts differed in one factor only. Brands were anonymized. Document order was swapped so content would not be confused with position bias. They never called a search engine. Retrieval was off.

Four factors were unanimous across all six models, with a huge effect (odds ratio above 100, which in this design means a near-deterministic preference in a two-page contest, not a 100x lift on Google):

- the text is about something else (topic mismatch)
- the price is missing
- the date is 2019 against a 2026 date
- the document is second in the injected list

Completeness and trust cues help less. Formatting-only edits (a dense paragraph versus sections, information scattered through the page) had no consistent effect. "Recent date versus no date" was not a consensus result either. What was consensus is a recent date against an old date [15].

The practical reading, with the lab limit in front: price and date do not get a page into the index. They break a tie between pages that were already retrieved. A service page with no price ("contact us") loses, in this test, to the page that states the number. A page dated 2019 loses to a page whose facts are from 2026. Stamping today's date on stale copy is not the factor they measured. They changed the dated content.

For a services firm, the page that clears this gate answers out loud, in the first paragraph:

- what is delivered (a WhatsApp agent that books appointments, an n8n automation, an app that outgrew Lovable)
- for whom (a clinic group, a hotel group, a B2B wholesaler)
- how long (most automations live in 2 to 4 weeks, a full project in 4 to 8 weeks)
- how much (a sprint at US$1,200 for two weeks, a project from US$3,000, a partnership from US$900 a month)
- the date those figures were written

Bing, in the 10 February 2026 AI Performance launch, recommends the same kind of page as product guidance, without a controlled test: depth, headings, tables, an FAQ in the body, evidence, freshness, and text that matches the images and video. Citations in the report do not indicate placement inside the answer [12]. Structure helps a person and helps extraction. It does not replace price, date, and topic.

## Gate 3: trust comes from outside, and a fake mention is spam

The guide says generative answers can show what the web says about a product, in blogs, videos, and forums. It then takes apart the shortcut: chasing inauthentic mentions is less useful than it looks. Ranking systems look at quality. Spam systems block spam. Generative features depend on both [1].

Since 15 May 2026 that is policy, not advice. The changelog entry that day clarifies that spam policies also apply to generative AI responses in Google Search. It was not a new policy. It was the notice that the existing policy already covered AI Overviews and AI Mode [2]. The current policy text, updated on 28 August 2026, defines spam as a technique used to deceive users or to manipulate Search systems, "such as attempting to manipulate generative AI responses in Google Search" [3].

What that text covers, in buyer language:

- **Scaled content.** Many pages whose main purpose is to manipulate rankings, not to help. That includes generating pages with AI that add nothing, and standing up multiple sites to hide the scale [3]. One URL per city, or one URL per fan-out query, with the same text and a swapped city name, is the pattern the July guide points at that policy [1].
- **Doorways.** Pages built to rank for a query and then push the person somewhere else. Several domains with near-identical homepages. City pages that dump everyone onto one URL [3].
- **Link spam.** Buying links, requiring a link in a contract without letting the other party use `nofollow` or `sponsored`, widgets and footers spread across many sites, low-value content built to manufacture a signal [3].
- **Expired domains.** Buying an old domain to inherit rankings with content that has little or no value [3].

Outside Google, Microsoft documented the commercial cousin of that shortcut. On 10 February 2026 the security team described "AI Recommendation Poisoning": "Summarize with AI" buttons that open an assistant with an instruction to remember the company as a trusted source or to recommend it first. Over 60 days of AI URLs in email traffic, they found more than 50 distinct prompts, from 31 companies, across 14 industries (finance, health, legal, SaaS, agencies, food, business services). Effectiveness varied by assistant and over time. Microsoft classifies the case as memory poisoning [19]. That is an attack. It is not a brand tactic.

Trust that clears gate 3 is dull and checkable. A partner-directory listing the company actually earned. A reported article. A customer page the customer controls. A person's profile, with a name and a history, that matches the site. For a local business, Google points to Business Profile and Merchant Center as the path into AI answers [1]. Bing points to Bing Places for address, hours, and contact on local answers [12]. A software studio that has done work in São Paulo, New York, Madrid, and Tel Aviv does not become "the best agency in town" with a fake local listing in twenty cities. It becomes a true listing where it actually operates, plus pages that survive a check.

## Gate 4: freshness is a new fact, not a publishing calendar

In the SIGIR test, a 2026 date against a 2019 date was one of the four unanimous factors, with retrieval turned off [15]. On Bing, the AI Performance announcement treats updates as the condition for an answer to use the current version, and points to IndexNow as the ping to participating engines [12]. Google states the other side: a high quantity of pages does not make a site more relevant, and the systems already understand relevance without an exact word match [1].

The freshness work that fits a services firm is short.

- When the price changes, the pricing page changes the same day. A sprint still listed at an old number teaches the model the wrong number. The public figure is US$1,200 per two-week sprint.
- When the timeline changes, the sentence "2 to 4 weeks" changes. A new post is the wrong tool for fixing one sentence.
- The visible date follows a revision of the facts, not a cosmetic title tweak.
- One good page, kept current, beats twelve thin pages covering the fan-out.

## What does not move citations

The 10 July 2026 guide has a mythbusting section. Four items match what agencies still sell [1]. The 15 June 2026 changelog had already added the llms.txt note: the file is not needed for Google Search, and it does not change visibility or rankings in either direction [2].

| Tactic | What the source says | Note |
|---|---|---|
| llms.txt | Google Search does not use it. Keeping the file neither helps nor hurts on Google [1]. | Vendor study: 97% of valid files received zero requests in May 2026 [16]. Detail in the [llms.txt piece](/blog/does-llms-txt-work/). |
| Special AI schema | Not required for generative search. There is no special schema.org markup for it. Schema remains useful for classic rich results [1]. | Vendor, difference-in-differences: 1,885 pages that gained JSON-LD between August 2025 and March 2026, against about 4,000 controls. AI Mode up 2.4% and ChatGPT up 2.2%, indistinguishable from zero. AI Overviews down 4.6%, a small effect. The pages were already heavily cited. The test does not speak for pages that are not cited yet [17]. |
| Chunking | There is no requirement to break the page into tiny pieces. Google says it can understand several topics on one page [1]. | In the SIGIR lab, formatting alone had no consistent effect [15]. |
| Inauthentic mentions | The guide says to ignore them. Spam policy, since 15 May 2026, covers manipulating an AI answer [1] [2] [3]. | A "remember us" button was documented by Microsoft as memory poisoning [19]. |
| Rewriting only for the AI | Not required. The systems understand synonyms and intent [1]. | A page written for a person, with a checkable fact, is what the guide asks for. |

Google's FAQ rich result was deprecated in the 8 May 2026 changelog entry [2]. An FAQ in the body of the page is still text. An FAQ badge in classic search is not a citation lever.

## What has counted as spam since 15 May 2026

That date did not invent a new penalty with GEO in the name. It made explicit that manipulating a generative AI response is Search spam [2] [3]. In practice, these jobs now have a name in the policy if the main purpose is manipulation:

- a network of satellite sites citing each other
- dozens of near-duplicate pages, one per fan-out variant or per city
- hidden text with an instruction to the model
- a link required by contract, without `rel="sponsored"` or `rel="nofollow"`
- an expired domain stuffed with commercial pages that add nothing

Detection is automated and, when needed, human review that can end in a manual action. The site can rank worse or drop out of results [3]. "An agency published it" is not a defense written into the policy.

## How to measure this without fooling yourself

Position inside one AI answer is not a metric. Presence, across many runs, is what the official instruments and the consistency test can support.

Search Console has a generative AI performance report with impressions of links in AI Overviews and AI Mode: pages, countries, dates (Pacific Time), and devices. Since 31 August 2026 the report is available to sites worldwide, if there are enough impressions. The help page, as opened, has no columns for clicks, CTR, position, or query. The same data also sits inside the Web performance report. Adding one on top of the other counts the same impression twice [11].

Bing Webmaster Tools has shown citations since 10 February 2026: total citations, average cited pages per day, a sample of grounding queries, citations by URL, and the time series. The announcement says, in more than one place, that this does not indicate placement, ranking, or the role of the page inside a single answer [12]. On 16 June 2026, Intents, Topics, Citation Share, and Compare entered a global preview. Citation Share is the percentage of citations for that grounding query that went to your site. Microsoft writes that it is not a ranking, not a traffic share, does not reveal a competitor's domain, and does not assign a quality score [13].

The consistency test is Rand Fishkin's, at SparkToro, with collection in November and December 2025. About 600 volunteers ran 12 prompts on ChatGPT, Claude, and Google's AI, 2,961 times in total. The chance that the same brand list appears in two answers was under 1 in 100. The same list in the same order was about 1 in 1,000. What repeated was the set of brands that show up often (a presence rate), not the order. The data partner works at Gumshoe, an AI tracking company. The write-up treats position as a metric with no foundation, and presence across many runs as the measure that remains [18].

What is not a measurement:

- a screenshot of one chat
- "we came in second" in a single answer
- a third-party tool that promises an internal Google metric (the guide says no third party has access to those systems) [1]
- adding generative AI impressions on top of Web performance [11]

An honest cadence, once Search Console and Bing are verified: look at both reports every week, and run a fixed set of real customer questions many times a month, recording whether the brand appeared. The trend of that set is the result. One day's swing is not.

## Common mistakes

- Blocking GPTBot and concluding the site "left ChatGPT." The search switch is OAI-SearchBot [4].
- Allowing the bot in robots.txt and blocking the IP at the edge. OpenAI conditions eligibility on both [5].
- Treating `Google-Extended` as the AI Mode switch [9].
- Publishing twenty cities with no operation, to "cover the fan-out" [1] [3].
- Hiding the price and hoping the model invents a flattering number. In the lab, a missing price was disqualifying [15].
- Measuring a screenshot and calling it a rank [18].
- Buying an llms.txt generator and calling that the program [1] [16].

## What not to do

Do not buy mentions. Do not build a domain network. Do not hide an instruction to the model. Do not publish a page per synonym. Do not rewrite the site "in the AI's tone." Do not swap useful copy for FAQ markup aimed at a rich result. Do not promise, and do not accept a promise of, first place in ChatGPT.

If the true page does not exist yet, the work is to write one: the service, the price, the timeline, the date, an anonymized example of work already shipped. A multi-tenant B2B SaaS for gyms that started in Lovable and now runs in production with paying customers. WhatsApp agents in production for a clinic group, for hospitality and B2B retail, for an adventure park, for a property manager, and for a visa advisory. That is gate 2 material. Ten variations of the same paragraph are not.

## When not to hire samambai

Do not hire samambai if the ask is any of these:

- a guarantee of position, a citation percentage, or "first in ChatGPT in 30 days"
- a site network, a paid mention, or a button that asks the model to memorize the brand
- an llms.txt file, a schema install, or "chunking" as the whole deliverable
- paid media management (not part of what samambai sells)
- a result next week from a domain Google and Bing have not indexed, with no new page and no willingness to wait for a crawl

The work fits when there is a real service, a page that can state the price, and patience to measure presence. The first conversation is 30 minutes, with no commitment. The public formats are on [pricing](/pricing/): a sprint at US$1,200 for two weeks, cancel any time; a project from US$3,000, with scope, timeline, and price fixed, and a delivery every sprint; a partnership from US$900 a month. The GEO service is described on the [GEO page](/geo/). Why the llms.txt file is not this work is in the [companion piece](/blog/does-llms-txt-work/).

The criterion is the same one used on every engagement: diagnosis, strategy, execution, and evolution, in two-week sprints, with a human review. AI agents sit inside the build, the tests, and a separate audit. They do not replace someone reading the page before it goes live.
