How Do AI Engines Decide Which Brands to Cite?
Ask ChatGPT which vendor to use in almost any category and it names three to five companies. Those names decide who gets the enquiry. This article covers what is actually known about how they get chosen, from published research and the platforms' own documentation. We run citation measurements across these engines every day, and most of what circulates about "ranking in AI" is guesswork, so it is worth separating what is verified from what is sold.
Two ways an AI answers
An AI answers from one of two places. Either from its training memory, which is frozen months in the past, or from live retrieval, where it searches the web mid-answer and reads what it finds. Commercial questions increasingly trigger retrieval: ChatGPT runs live search through its search index when it judges the question needs current information, Claude searches when the question needs verifiable facts, and Google's AI features retrieve by design.
Retrieval is where citation happens. When the engine reads pages to build an answer, the pages it leans on become the sources it names.
The fan-out
Google has documented this part itself. Its Search Central documentation says both AI Overviews and AI Mode may use a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources," to develop a response. One buyer question becomes many sub-queries, each pulling its own results, and the answer is synthesised across all of them. This is why AI answers cite a wider and stranger set of pages than the top ten blue links. Your brand can enter the answer through a sub-query you never optimised for.
What moves citations, per the research
The founding study here is GEO: Generative Engine Optimization (Aggarwal et al., presented at KDD 2024), which tested nine optimisation tactics across 10,000 queries. Three of them moved visibility, and all three landed in the same band, roughly 30 to 40% depending on which metric you measure:
- Adding quotations from credible sources.
- Adding statistics.
- Citing sources.
Which of the three comes out strongest shifts with the metric, so they are better read as one family of tactics than as a ranking. Keyword stuffing did nothing, and measured slightly negative on Perplexity. Persuasive "authoritative" tone also did nothing.
The pattern is consistent: engines reward content that looks like evidence, and ignore content that looks like marketing.
Being cited is not being recommended
Lily Ray's 2026 study ran 100 B2B software queries, 80 of which triggered a Google AI Overview. When brands published self-promotional "best software" listicles, those Overviews cited the brand's pages 323 times, but in 224 of those cases cited the page without recommending the brand anywhere in the answer. The engine mined the listicle for candidates and sometimes recommended the competitors listed in it. Publishing a comparison does not mean winning it.
The engines do not agree with each other
Kevin Indig analysed 3.7 million citations across ChatGPT, Perplexity and Google AI Overviews. Within a 20,000-prompt sample run identically on all three, only 2.37% of cited URLs appeared in every engine, and 91% appeared in exactly one. Each engine is closer to a separate retrieval pool than a shared ranking, which means visibility has to be measured per engine. Being everywhere in ChatGPT says nothing about Perplexity.
llms.txt
llms.txt is a proposed file that describes your site to AI systems. We deploy it for clients, so here is the straight version: as of mid-2026, no major engine has publicly confirmed reading it. OpenAI's crawler documentation covers robots.txt and never mentions it, Perplexity's likewise, and Google's John Mueller has said plainly that Google does not use it, adding that none of the AI services have said they do. The confusing part is that Google and Anthropic both publish llms.txt files for their own documentation sites. Publishing one and reading yours are different things.
What gets a brand into ChatGPT or Claude answers is being retrievable through the search indexes those engines query mid-answer, which is ordinary crawlability, plus content worth citing. We treat llms.txt as a few minutes of work with possible future value and zero downside, and anyone selling it as a ranking signal, for any engine, is misinforming you.
What this means in practice
Citability is mostly evidence-shape: pages with concrete numbers, quotable claims, named sources and clean structure get pulled into answers. Third-party corroboration matters more than your own site, since engines lean heavily on community and review platforms. And because engines disagree, measurement has to cover each one separately, repeatedly, since answers also change run to run.
Sources
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
- Google Search Central, AI features and your website, updated June 2026
- Lily Ray, 100-query AI Overviews study, via Search Engine Land, June 2026
- Kevin Indig, The Consensus Gap, Growth Memo, May 2026
- John Mueller on llms.txt, via Search Engine Roundtable, January 2026
- Ahrefs, AI Overviews change 70% of the time, November 2025
- OpenAI Help Center, ChatGPT Search
- Anthropic, Web search tool documentation
Questions, answered
Does schema markup help AI citations?
It helps machines parse who you are and what you claim, and it is standard hygiene for AI readability. No published study shows schema alone moving citation rates the way quotations and statistics do.
Can you pay to be cited?
No. None of the major engines sell citation placement in organic answers today.
How stable are AI citations?
Not very. Ahrefs measured a 70% chance an AI Overview changes between observations, with only 54.5% of cited URLs overlapping between refreshes. Visibility is a distribution over time, never a screenshot.
See where you stand.
Start with a free audit. We will show you exactly where you are cited across AI engines, and where you are not.
Get a free audit