How AI Assistants Choose Which Sources to Cite
There is no single AI search algorithm. Each platform retrieves and cites differently, and those differences decide what is worth commissioning for a client and what is wasted effort. Where a number could not be traced to a source, it has been removed.
ChatGPT, Perplexity, Google AI Overviews, Claude and Gemini each use a distinct index, retrieval method and reranking model. There is no unified AI search algorithm, and any vendor who implies otherwise is selling you a simplification.
For an agency the practical consequence is about commissioning. Some content decisions pay off across every platform. Others are specific to one. Knowing which is which is the difference between a content plan and a scattergun.
A note on the numbers: this is the page where the industry’s folklore is thickest. Every figure below carries a footnote to a page you can read; the many “citation boost” multipliers that circulate without a traceable study have been left out.
As of September 2026, ChatGPT holds an estimated 51.5% of US generative AI chatbot usage, Google’s Gemini 27.6%, Claude 10.2%, Grok 2.8%, Perplexity 2.0% and Microsoft Copilot 1.3%.1 Optimising for one platform means being absent from a large part of the channel, and the second-placed platform has been the fastest-growing.
| Platform | Share (Sept 2026)1 | Primary index | Defining behaviour |
|---|---|---|---|
| ChatGPT | 51.5% | Bing, increasingly Google2 | Live retrieval, filters for readable sources |
| Gemini | 27.6% | Knowledge Graph and entity records | |
| Claude | 10.2% | Brave Search3 | Direct answers, inline citations |
| Grok | 2.8% | X and web search | Real-time social signal |
| Perplexity | 2.0% | Own crawler | Reranking with heavy freshness weight |
| Microsoft Copilot | 1.3% | Bing | Deep Bing integration |
ChatGPT
ChatGPT answers by running live searches and applying its own scoring to the results. Which index it leans on has been shifting.
Profound analysed a sample of its 240 million ChatGPT citations and found the share of ChatGPT’s cited sources that also appear in Google’s results rose from 12% in April 2025 to 33% in July 2025, while overlap with Bing’s results fell from 26% to 8% over the same three months.2 Google indexation now matters for ChatGPT visibility in a way it did not in 2024.
What can be observed from the outside is that ChatGPT favours sources that present verifiable, linkable, readable information, and that paywalled and login-gated material does not surface.
What this changes about commissioning. Anything substantive that lives behind a gate is invisible here. If a client’s best material is all gated whitepapers, the programme’s first act is negotiating an ungated tier, an argument that is far easier to win with this mechanism explained than as a general principle about openness.
Perplexity
Perplexity fetches content at query time rather than serving from a static pre-built index, and cites a short list of sources inline. Well-structured material can be cited soon after publication, which is dramatically faster feedback than conventional search offers.
Freshness is its most visible signal. Explicit dates in the prose (“as of September 2026”) beat relative language (“recently”). We have not found a study that puts a defensible number on how much freshness matters here, so we do not print one.
What this changes about commissioning. Perplexity is the strongest argument for a refresh budget. A plan that publishes and moves on underperforms one that publishes less and revisits more. If a client wants early evidence that a programme is doing anything at all, this is the platform where it shows first.
Google AI Overviews
AI Overviews select sources through query fan-out: decomposing one user question into many simultaneous sub-queries across Google Search, the Knowledge Graph, News, Shopping and other sources, then synthesising from the strongest results.
Surfer’s study of 173,902 URLs across 10,000 keywords found a 0.77 correlation between the number of fan-out sub-queries a page ranks for and its probability of being cited; pages ranking for fan-out queries were 161% more likely to be cited than pages ranking only for the main query; and 68% of the pages cited were not in the top ten organic results for the query itself.4 Content covering a topic and its adjacent sub-topics comprehensively performs far better than content targeting a single keyword.
Being cited in the overview is worth having: Seer Interactive found brands cited inside an AI Overview earn about 35% more organic clicks than uncited brands on the same query.5
What this changes about commissioning. This is the platform that rewards depth over breadth. One thoroughly worked topic cluster beats twenty thin pages across twenty topics. Named authorship with real credentials is not decoration here; Google’s own quality guidance treats it as a signal, which is worth remembering when a client asks whether bylines matter.
Claude
Claude retrieves through Brave Search: Anthropic added Brave Search to its list of subprocessors on 19 March 2025, the day before it launched web search in Claude.3 It applies its own filtering before generating a response with inline citations.
Its selection priorities are the most editorially opinionated of the five. It favours content that answers a question directly rather than hedging; it needs verifiable, specific anchors - named entities, statistics, dates - to generate a cited response; and it discounts promotional language. Superlatives without data behind them (“industry-leading”, “best-in-class”) do not give it anything to cite.
What this changes about commissioning. Claude is the clearest argument against the house style most B2B marketing sites are written in. If a client’s pages open with three sentences of positioning before reaching a fact, that is a liability rather than a matter of taste. Lead with the answer inside the first 60 words of a section.
Gemini
Gemini draws on Google’s full ecosystem - organic rankings, Knowledge Graph entity records, and Google’s own data layers - making it the platform most tightly coupled to a client’s existing Google presence. It has also been the fastest-growing of the assistants, rising to 27.6% of US chatbot usage by September 2026.1
Entity clarity is the lever. Google’s models validate claims against the Knowledge Graph, and inconsistent inputs - conflicting descriptions, mismatched names, contradictory facts across a client’s own properties - introduce ambiguity that reduces recommendation frequency.
What this changes about commissioning. Gemini is where existing SEO investment converts. If a client already ranks, the incremental work is entity discipline and answer-shaped content, not a rebuild. That is a useful thing to be able to say to a client who fears they are being sold a second, parallel programme alongside the one they already fund.
The comparison, in one table
| Factor | ChatGPT | Perplexity | Google AIO | Claude | Gemini |
|---|---|---|---|---|---|
| Primary index | Bing, increasingly Google | Own crawler | Google + Knowledge Graph | Brave Search | Google ecosystem |
| Retrieval | Live search | On-demand crawl | Query fan-out | Brave results | Google index + entity records |
| Freshness weight | Moderate | Very high | High | High | Moderate-high |
| Structured data | Helps | FAQPage schema | Helps | JSON-LD + clarity | Article + entity schema |
| Gated content | Skipped | Skipped | Skipped | Skipped | Skipped |
| Strongest lever | Google and Bing indexation, ungated prose | Freshness, refresh cadence | Topical depth, fan-out coverage | Direct answers, no puffery | Entity consistency, Google presence |
What pays across all five
Five commissioning decisions improve visibility on every platform at once. These are the defensible core of a content plan.
-
Let the crawlers in. GPTBot, ClaudeBot, PerplexityBot, Bingbot, Googlebot. Blocking any one removes that platform entirely. Check the CDN configuration as well as
robots.txt: the block is often not where you expect it. -
Write direct answers. Every question-shaped heading followed by a 40-80 word answer that stands alone without surrounding context. This single format satisfies the extraction behaviour of all five platforms, and it is a writing brief rather than a technical spec.
-
Implement schema properly. FAQPage, Article and Organization JSON-LD give every platform cleaner entity data to work from. We do not print a citation multiplier for schema because we could not trace one to a study.
-
Keep the entity consistent. One name, one description, one set of core facts, everywhere. This is a governance job more than a creative one, and it is usually the fastest thing to fix across a client’s estate.
-
Attribute everything. The Princeton GEO study (KDD 2024) found adding statistics raised visibility in AI-generated answers by 33% and adding quotations by 41%.6 Every substantive claim should carry a source, a figure and a year. This is the largest single gap between typical agency content and content that actually gets cited, and it is the rule this page is written to.
Frequently asked questions
Does Google SEO get a client into ChatGPT? Increasingly. The overlap between ChatGPT’s citations and Google’s results rose from 12% to 33% between April and July 2025,2 so rankings help more than they used to. But ChatGPT independently filters for readability and verifiability. Ranking on Google helps; it does not guarantee anything.
What is query fan-out? Google AI Overviews’ technique of splitting one query into many sub-queries across its data sources. A study of 173,902 URLs found a 0.77 correlation between the number of fan-out queries a page ranks for and its citation probability.4 It rewards comprehensive coverage of a topic and its neighbours.
How does Perplexity decide what to cite? It fetches at query time and weights freshness heavily. Recently updated, plainly structured, explicitly dated content is what it reaches for.
Which index does Claude use? Brave Search.3 Brave indexation is the technical prerequisite; direct, unpromotional writing is the editorial one.
Why does each platform need separate consideration? Different indexes, different rerankers, different content requirements. A client strong in one may be absent from another. With ChatGPT at roughly half of chatbot usage and Gemini at more than a quarter,1 ignoring either means missing a large part of the channel.
What structured data matters most? FAQPage, Article, HowTo and Organization are the types every platform reads. Treat them as hygiene, not as a lever with a number attached.
Key takeaways
- Five platforms, five indexes: Bing and Google (ChatGPT), its own crawler (Perplexity), Google’s ecosystem (Gemini and AI Overviews), Brave (Claude).
- No single action covers all five, but five come close: crawler access, direct answer blocks, schema, entity consistency, attributed statistics.
- Perplexity rewards freshness. Budget for refreshes or accept decline.
- Google AI Overviews reward depth and adjacency: pages ranking for fan-out queries are 161% more likely to be cited.4
- Claude discounts promotional language, which makes house style a measurable factor rather than a stylistic preference.
- Gemini converts existing Google investment, and it is the fastest-growing assistant.
- The Princeton findings - statistics +33%, quotations +41%6 - are editorial instructions, not technical ones. They are also the hardest thing to hold to at volume.
Published February 2026, figures re-checked September 2026. Found by AI monitors how leading AI assistants answer buying questions, produces the content that addresses the gaps, and reports on what the published work earns - delivered through white-label agency partners. See how the partner programme works.
Sources
Every figure above links to the page it was taken from. Checked on 10 September 2026.
Footnotes
-
First Page Sage, “Top Generative AI Chatbots by Market Share, September 2026”, estimates based on monthly active users across US web and mobile. ↩ ↩2 ↩3 ↩4
-
Profound, “AI Search Shift: ChatGPT’s growing alignment with Google’s index”, sample from a dataset of 240 million ChatGPT citations, April to July 2025, published 6 August 2025. ↩ ↩2 ↩3
-
Anthropic’s subprocessor list, documented by Simon Willison, “Anthropic Trust Center: Brave Search added as a subprocessor”, 21 March 2025. ↩ ↩2 ↩3
-
Surfer, “Query Fan-Out Impact”, 173,902 URLs across 10,000 keywords, 2025. ↩ ↩2 ↩3
-
Seer Interactive, “AIO Impact on Google CTR: September 2025 Update”, June 2024 to September 2025. ↩
-
Aggarwal et al., “GEO: Generative Engine Optimization”, Proceedings of the 30th ACM SIGKDD Conference, 2024, Table 1. ↩ ↩2
What to do with this
The gaps are the brief. We write what fills them.
Monitoring shows which buying questions the leading AI assistants answer with somebody else’s name. Those questions become the month’s articles, answer blocks, fact sections, FAQs and refreshes - written, published and measured under our partners’ brands.