How AI Assistants Choose Which Sources to Cite
There is no single AI search algorithm. Each platform retrieves, reranks and cites differently - and those differences decide what is worth commissioning for a client and what is wasted effort.
ChatGPT, Perplexity, Google AI Overviews, Claude and Gemini each use a distinct index, retrieval method and reranking model. There is no unified AI search algorithm, and any vendor who implies otherwise is selling you a simplification.
For an agency the practical consequence is about commissioning. Some content decisions pay off across every platform. Others are specific to one. Knowing which is which is the difference between a content plan and a scattergun.
As of February 2026, ChatGPT holds 60.5% of AI search market share, Microsoft Copilot 14.3%, Gemini 13.5%, Perplexity 6.2% and Claude 3.2% (First Page Sage, February 2026). Optimising for one platform means being absent from most of the channel.
| Platform | Share (Feb 2026) | Primary index | Defining behaviour |
|---|---|---|---|
| ChatGPT | 60.5% | Bing + Google (fallback) | Live retrieval, filters for readable sources |
| Microsoft Copilot | 14.3% | Bing | Deep Bing integration |
| Gemini | 13.5% | Knowledge Graph and entity records | |
| Perplexity | 6.2% | Own Sonar crawler | Reranking with heavy freshness weight |
| Claude | 3.2% | Brave Search | Precision filtering, inline citations |
Source: First Page Sage, February 2026
ChatGPT
ChatGPT answers by running live searches against Bing’s index - increasingly against Google’s as a fallback - and applying its own scoring to the top 20–30 results.
Its alignment with Google Search results rose from 12% to 33% between April and July 2025, while Bing alignment dropped from 26% to 8% (Profound, 2025). Google indexation now matters for ChatGPT visibility in a way it did not in 2024, but Bing indexation remains the more reliable lever.
The selection sequence, as far as it can be observed: a Bing search across roughly 20–30 results, a shortlist of five to eight promising sources, then three to five finalists that present verifiable, linkable, readable information. Paywalled and login-gated material is filtered out at this stage.
What this changes about commissioning. Anything substantive that lives behind a gate is invisible here. If a client’s best material is all gated whitepapers, the programme’s first act is negotiating an ungated tier - an argument that is far easier to win with this mechanism explained than as a general principle about openness.
Perplexity
Perplexity uses retrieval-augmented generation with real-time crawling and an XGBoost reranker, visiting roughly ten pages per query and citing three or four.
Rather than a static pre-built index, its Sonar model fetches content at query time. Well-structured material can be cited within hours of publication - dramatically faster feedback than conventional search offers.
Freshness is its dominant signal. An article published or updated two hours ago is cited 38% more often than an identical article with last month’s dateline (Growth Memo, 2026). Content decay begins two to three days after publication without updates - a far more aggressive recency curve than any other platform applies.
Domain authority accounts for only about 15% of its ranking. A well-researched piece on a newer domain can outrank an established source if it is more current, better structured and more semantically complete.
What this changes about commissioning. Perplexity is the strongest argument for a refresh budget. A plan that publishes and moves on underperforms one that publishes less and revisits more. Explicit dates in the prose (“as of February 2026”) beat relative language (“recently”). If a client wants early evidence that a programme is doing anything at all, this is the platform where it shows first.
Google AI Overviews
AI Overviews select sources through query fan-out - decomposing one user question into many simultaneous sub-queries across Google Search, the Knowledge Graph, News, Shopping and other sources, then synthesising from the strongest results.
A SurferSEO study of 173,902 URLs found a 0.77 correlation between the number of fan-out sub-queries a page ranks for and its probability of being cited. Content covering a topic and its adjacent sub-topics comprehensively performs far better than content targeting a single keyword.
Reported ranking factors include semantic completeness, E-E-A-T signals (Google’s experience-expertise-authoritativeness-trust standard), entity density (pages with 15+ recognised named entities per 1,000 words show a 4.8x citation boost), JSON-LD structured data (+73% selection rate), and freshness of cited statistics (+89%).
What this changes about commissioning. This is the platform that rewards depth over breadth. One thoroughly worked topic cluster beats twenty thin pages across twenty topics - citation rates for clusters rise from 12% to 41% against isolated pages in B2B studies. Named authorship with real credentials is not decoration here; it is a ranking input, which is worth remembering when a client asks whether bylines matter.
Claude
Claude retrieves through Brave Search, taking roughly the top ten results and applying its own filtering before generating a response with inline citations. Anthropic listed Brave Search as a subprocessor in early 2025 (TechCrunch, March 2025).
Its selection priorities are the most editorially opinionated of the five. It favours content that answers a question directly rather than hedging; it needs verifiable, specific anchors - named entities, statistics, dates - to generate a cited response; it weights year-specific phrasing; and it actively discounts promotional language. Superlatives without data behind them (“industry-leading”, “best-in-class”) reduce source confidence.
What this changes about commissioning. Claude is the clearest argument against the house style most B2B marketing sites are written in. If a client’s pages open with three sentences of positioning before reaching a fact, that is a measurable liability rather than a matter of taste. Lead with the answer inside the first 60 words of a section.
Gemini
Gemini draws on Google’s full ecosystem - organic rankings, Knowledge Graph entity records, and Google’s own data layers - making it the platform most tightly coupled to a client’s existing Google presence.
Sites in Google’s top ten are 3x more likely to be cited in Gemini responses than those beyond position 20 (MAK Digital Design, 2026). Conventional Google signals carry over more directly here than to any other assistant.
Entity clarity is the second lever. Google’s models validate claims against the Knowledge Graph, and inconsistent inputs - conflicting descriptions, mismatched names, contradictory facts across a client’s own properties - introduce ambiguity that reduces recommendation frequency.
What this changes about commissioning. Gemini is where existing SEO investment converts. If a client already ranks, the incremental work is entity discipline and answer-shaped content, not a rebuild. That is a useful thing to be able to say to a client who fears they are being sold a second, parallel programme alongside the one they already fund.
The comparison, in one table
| Factor | ChatGPT | Perplexity | Google AIO | Claude | Gemini |
|---|---|---|---|---|---|
| Primary index | Bing + Google | Own Sonar crawler | Google + Knowledge Graph | Brave Search | Google ecosystem |
| Retrieval | Live Bing search | On-demand crawl | Query fan-out | Brave top ~10 | Google index + entity records |
| Freshness weight | Moderate | Very high (2–3 day decay) | High | High | Moderate–high |
| Reranking | Internal scoring | XGBoost reranker | E-E-A-T + entity density | Internal filtering | Google ranking signals |
| Structured data | Helps | FAQPage schema | +73% with schema | JSON-LD + clarity | Article + entity schema |
| Gated content | Skipped | Skipped | Skipped | Skipped | Skipped |
| Strongest lever | Bing indexation, ungated prose | Freshness, refresh cadence | Topical depth, fan-out coverage | Direct answers, no puffery | Entity consistency, Google presence |
What pays across all five
Five commissioning decisions improve visibility on every platform at once. These are the defensible core of a content plan.
-
Let the crawlers in. GPTBot, ClaudeBot, PerplexityBot, Bingbot, Googlebot. Blocking any one removes that platform entirely. Check the CDN configuration as well as
robots.txt- the block is often not where you expect it. -
Write direct answers. Every question-shaped heading followed by a 40–80 word answer that stands alone without surrounding context. This single format satisfies the extraction behaviour of all five platforms, and it is a writing brief rather than a technical spec.
-
Implement schema properly. FAQPage, Article and Organization JSON-LD raise citation rates across Google AI Overviews (+73%), Perplexity and Gemini.
-
Keep the entity consistent. One name, one description, one set of core facts, everywhere. This is a governance job more than a creative one, and it is usually the fastest thing to fix across a client’s estate.
-
Attribute everything. The Princeton GEO study (KDD 2024) found adding statistics raises AI visibility by 33% and adding quotations by 41%. Every substantive claim should carry a source, a figure and a year. This is the largest single gap between typical agency content and content that actually gets cited.
Frequently asked questions
Does Google SEO get a client into ChatGPT? Partially. ChatGPT’s alignment with Google’s index rose from 12% to 33% during 2025 (Profound, 2025), so rankings help more than they used to. But ChatGPT independently filters for readability and verifiability. Ranking on Google helps; it does not guarantee anything.
What is query fan-out? Google AI Overviews’ technique of splitting one query into many sub-queries across its data sources. A study of 173,902 URLs found a 0.77 correlation between the number of fan-out queries a page ranks for and its citation probability. It rewards comprehensive coverage of a topic and its neighbours.
How does Perplexity decide what to cite? It visits around ten pages per query and cites three or four, scoring candidates on semantic depth, trust, freshness and engagement. Recently updated content is cited 38% more often than identical older content (Growth Memo, 2026).
Which index does Claude use? Brave Search. Brave indexation is the technical prerequisite; direct, unpromotional writing is the editorial one.
Why does each platform need separate consideration? Different indexes, different rerankers, different content requirements. A client strong in one may be absent from another. With ChatGPT at 60.5% share and Gemini at 13.5%, ignoring either means missing most of the channel.
What structured data matters most? JSON-LD raises citation selection rates by around 73% (aimodeboost.com, 2025). FAQPage, Article, HowTo and Organization are the highest-impact types.
Key takeaways
- Five platforms, five indexes: Bing plus Google (ChatGPT), Sonar (Perplexity), Google’s ecosystem (Gemini and AI Overviews), Brave (Claude).
- No single action covers all five, but five come close: crawler access, direct answer blocks, schema, entity consistency, attributed statistics.
- Perplexity decays content within days. Budget for refreshes or accept decline.
- Google AI Overviews reward depth and adjacency. Clusters beat isolated pages by a wide margin.
- Claude penalises promotional language, which makes house style a measurable factor rather than a stylistic preference.
- Gemini converts existing Google investment. It is the easiest win for clients who already rank.
- The Princeton findings - statistics +33%, quotations +41% - are editorial instructions, not technical ones. They are also the hardest thing to hold to at volume.
Published February 2026. Found by AI monitors how leading AI assistants answer buying questions, produces the content that addresses the gaps, and reports on what the published work earns - delivered through white-label agency partners. See how the partner programme works.
Sources: Seer Interactive (2024), Joe Youngblood (2025), Profound (2025), TechCrunch (March 2025), Growth Memo - State of AI Search Optimization 2026, Princeton GEO Study (KDD 2024), First Page Sage (February 2026), SurferSEO study (173,902 URLs), aimodeboost.com (2025), MAK Digital Design (2026).
What to do with this
The gaps are the brief. We write what fills them.
Monitoring shows which buying questions the leading AI assistants answer with somebody else’s name. Those questions become the month’s articles, answer blocks, fact sections, FAQs and refreshes - written, published and measured under our partners’ brands.