How AI Answer Engines Choose Sources: The Evidence
A graded review of the studies and vendor documentation on how ChatGPT, Google AI Overviews, Perplexity and Copilot choose the pages they cite.
- Author
- WavX Editorial Team
- Published
- 2026-10-02T09:00:00.000Z
- Updated
- 2026-10-02T09:00:00.000Z
- Organisation
- WavX Solutions
- Telephone
- +919310079927
Description
All articles SEO GEO SEO Strategy Digital Marketing
How AI answer engines choose what to cite: the evidence, graded
WavX Editorial Team Engineering & delivery team, WavX Solutions
Published 2 October 2026 14 min read 3,367 words
Custom software, built from scratch · Building since 2022 · Gurgaon, Delhi NCR
Part of our AI Development guide AI Development Company Summarise with AI ChatGPT Claude Perplexity Google AI
AI answer engines search first and write second. ChatGPT, Google's AI Overviews and AI Mode, Perplexity, Copilot and Claude each run one or more searches, read the pages that come back, and write an answer that links to some of them. So the first condition is plain: the engine's search crawler must be able to fetch your page. Beyond that, no vendor publishes its selection rules, and the published research supports far fewer conclusions than most articles on this subject claim. This review lists each common claim, the source behind it, the date, and how strong the evidence is.
How the evidence is graded
Not all evidence is equal. A vendor's documentation tells you its rules but not its ranking. A correlation across thousands of sites tells you what cited pages have in common but not what caused the citation. Each claim below carries one of five grades.
Grade
What it means
What it can support
Controlled experiment
One thing was changed and the result compared with an unchanged baseline
A causal claim, within the conditions tested
Matched-control study
Pages that changed are compared with similar pages that did not, on live engines
A causal claim with caveats
Large observational dataset
Many pages or prompts were measured, with no intervention
A pattern or correlation, not a cause
Vendor documentation
The company that runs the engine states how its product works
The rules for eligibility. Says little about ranking
Small sample or anecdote
A few prompts, one site, or a report from a third party
A reason to test, nothing more
Most of the independent research comes from companies that sell search and AI visibility software: Ahrefs, Semrush and Profound. Their datasets are large and their methods are published, but they are not neutral parties. That is noted here once and applies throughout.
What the engines themselves document
Every engine documents how a page becomes eligible. None documents how it picks between eligible pages.
Engine
What its own documentation states
Source and date
Google AI Overviews and AI Mode
A page must be indexed and eligible to be shown with a snippet. There are "no additional requirements" and no special optimisation. Both features may use "query fan-out", issuing several related searches. Responses are grounded in pages retrieved by Google's core ranking systems
Google Search Central, AI features page (last updated 10 December 2025) and generative AI optimisation guide (last updated 10 July 2026)
ChatGPT search
OAI-SearchBot is the crawler used to surface sites in ChatGPT search. Sites that opt out of it "will not be shown in ChatGPT search answers", though they may still appear as navigational links. A robots.txt change takes about 24 hours to apply
OpenAI crawler documentation, read 2 October 2026
Claude
Claude-SearchBot indexes content for search, and Claude-User fetches pages when a user asks. Blocking either "may reduce your site's visibility" in user search results
Anthropic help centre, read 2 October 2026
Perplexity
PerplexityBot is "designed to surface and link websites in search results on Perplexity" and should be allowed in robots.txt for a site to appear. Perplexity-User fetches pages on a user's request and generally ignores robots.txt
Perplexity documentation, read 2 October 2026
Copilot and Bing
Copilot breaks pages into smaller pieces and assembles answers from several sources. Crawlability, metadata, internal linking and backlinks "remain essential". Microsoft adds that no practice ensures selection
Microsoft Advertising blog, 8 October 2025
Grade for all of the above: vendor documentation. It is reliable on eligibility and should be treated as the starting checklist. The crawler names and what each one does are collected in the AI crawler directory.
One more vendor fact matters. Since 2026 Google has offered a Search Console setting, the Search generative AI control, that includes or excludes a site from AI Overviews and AI Mode. Inclusion is the default. A site that has been excluded there cannot appear, whatever else it does.
The claims, graded
Claim
Evidence
The engine's search crawler must be able to fetch the page
Stated by every vendor
Google, OpenAI, Anthropic, Perplexity documentation
Vendor documentation. Treat as a requirement
Pages cited in AI Overviews usually rank in Google's top 10
76% of cited pages ranked in the top 10 in July 2025. An updated study found about 38% in March 2026
Ahrefs, 21 July 2025 and 2 March 2026
Large observational dataset. The overlap is real and has fallen
ChatGPT, Gemini and Copilot cite what Google ranks
Across 15,000 prompts, about 12% of cited links were in Google's top 10 for the same prompt. Perplexity was the exception at 28.6%
Ahrefs, 11 August 2025
Brands mentioned on other sites are mentioned more in AI Overviews
Across 75,000 brands, branded web mentions had the strongest correlation with AI Overview mentions (0.664)
Ahrefs, 26 May 2025
Large observational dataset. Correlation only
Adding statistics, quotations and cited sources raises visibility
30 to 40% relative improvement on one visibility measure, in a simulated engine
Aggarwal and others, KDD 2024 (arXiv 2311.09735)
Controlled experiment, not run on today's engines
"Best X" list articles are what ChatGPT cites for recommendations
Across 750 prompts, such lists were 43.8% of cited page types
Ahrefs, 4 December 2025
Observational, with manual categorisation
Each engine prefers different sources, and the mix shifts
Reddit and Wikipedia shares differ by engine and moved sharply within weeks
Profound, June 2025; Semrush, 10 November 2025
Large observational datasets
Schema markup gets pages cited
No meaningful uplift after 1,885 pages added JSON-LD, against 4,000 controls
Ahrefs, 11 May 2026
Matched-control study. The best evidence on this question
An llms.txt file gets pages cited
97% of published files received no requests in a month. Google states Search ignores the file
Ahrefs, 15 June 2026; Google, 10 July 2026
Large log study plus vendor documentation
An AI "ranking" can be tracked like a search ranking
The same prompt almost never returned the same list of brands twice
SparkToro and Gumshoe, 27 January 2026
Crowd experiment, not peer reviewed
The sections below explain what each of these does and does not show.
Retrieval comes before everything
Google describes its AI features as retrieval-augmented generation: its systems retrieve pages from the Search index using the core ranking systems, then generate a response from them, with links. It also describes query fan-out, where the model issues a set of related queries to gather more results. OpenAI, Anthropic and Perplexity describe separate crawlers for search and for user-requested fetches.
Two practical consequences follow, and both are well supported.
First, a page that a search crawler cannot fetch cannot be cited from search. Check robots.txt, any firewall or bot-protection rule, and whether the content is present in the HTML the server returns. Microsoft's guidance makes a related point: content hidden in tabs or expandable menus may not be rendered and can be skipped.
Second, the page competes on several queries, not one. Fan-out means the engine may retrieve your page for a sub-question the user never typed. Google's guide warns against reacting to this by publishing a page for every variation: doing so mainly to influence rankings or AI responses falls under its scaled content abuse policy.
How much does ranking in search matter?
Less than it did, and it depends on the engine.
For Google AI Overviews, Ahrefs looked at 1.9 million citations from 1 million AI Overviews in July 2025 and found 76.1% of cited pages ranking in the top 10 for the same query. It repeated the exercise on 863,000 keyword results pages and 4 million cited URLs, published on 2 March 2026. This time 37.9% of cited URLs were in the first ten results, 31.2% were in positions 11 to 100, and 31.0% were beyond 100. Ahrefs attributes the drop to greater reliance on fan-out queries. The two studies did not use identical methods (the first looked only at the three most visible citations), so the size of the fall should be read with care. The direction is consistent with Google's own description of fan-out.
For the assistants, the August 2025 Ahrefs study of 15,000 prompts found that on average 12% of links cited by ChatGPT, Gemini and Copilot appeared in Google's top 10 for the same prompt, and about 80% did not rank in Google for that prompt at all. Perplexity overlapped most, at 28.6%.
What this supports: ranking well in Google is still associated with being cited in AI Overviews, and it is a weak predictor for the assistants. What it does not support: the idea that search ranking no longer matters. The assistants still retrieve from search indexes, but for queries they write themselves.
Brand mentions: the strongest correlation, and only a correlation
Ahrefs' May 2025 study of 75,000 brands used Spearman correlation to compare several factors with how often a brand was mentioned in AI Overviews. Branded web mentions came first at 0.664, then branded anchor text (0.527) and branded search volume (0.392). Domain Rating was 0.326, the number of backlinks 0.218, and the number of pages on the site 0.17. The study also reports that 26% of brands had no AI Overview mentions at all.
The authors state the limit themselves: correlation is not causation. Well-known brands are mentioned more everywhere, including in AI answers. The sample was also restricted to domains with a Domain Rating above 40, so it describes established brands, not new ones.
Google's guide addresses the obvious shortcut. Seeking inauthentic mentions across the web "isn't as helpful as it might seem", because the AI features depend on the same ranking and spam systems as the rest of Search. Google's spam policies, updated on 28 August 2026, now name "attempting to manipulate generative AI responses" as an example of spam.
Statistics, quotations and cited sources: a controlled result with limits
The one controlled experiment in this field is the paper that coined the term generative engine optimisation, published at KDD 2024. The researchers built a benchmark of 10,000 queries and tested nine ways of rewriting a source page. Their engine retrieved the top five Google results for each query and generated an answer with GPT-3.5.
Three methods stood out. Adding citations from credible sources, adding quotations and adding statistics each improved a page's share of the answer by 30 to 40% on a position-adjusted word count measure, and by 15 to 30% on a subjective impression measure. Keyword stuffing did not help. In a smaller check on Perplexity, adding quotations improved the first measure by 22% and keyword stuffing performed 10% worse than the baseline.
One table in the paper deserves more attention than it gets. The gains went mostly to lower-ranked sources. Adding citations raised visibility by 115.1% for the source ranked fifth and lowered it by 30.3% for the source ranked first.
The limits are significant. The pages tested had already been retrieved, so the paper says nothing about getting retrieved in the first place. The main engine was built by the researchers on a model that is now several generations old. The measure is share of words in an answer, not clicks or enquiries. The authors themselves note that methods may need to change as engines evolve. It is fair to say that sourced, specific content was favoured in a controlled setting. It is not fair to quote "40% more visibility" as a result anyone should expect.
List articles and the recommendation question
For prompts that ask which product, tool or agency to use, Ahrefs analysed ChatGPT's responses to 750 prompts and categorised 26,283 source URLs. "Best X" list articles were 43.8% of all page types cited. Of 1,100 such lists with a clear date, 79.1% had been updated in 2025, and 35% of the cited lists sat on low-authority domains, many of which the author describes as questionable.
This is descriptive. It shows what ChatGPT drew on in late 2025 for one kind of prompt. The author does not claim that publishing a list causes recommendations, and notes that many recommendations come from multi-step conversations the study did not cover. A business reading this should also weigh Google's spam policy before publishing a list that ranks itself first.
Each engine cites differently, and the mix moves
Profound analysed 680 million citations between August 2024 and June 2025. Wikipedia was ChatGPT's most cited source at 7.8% of its citations. Reddit led on Perplexity at 6.6% and on Google AI Overviews at 2.2%.
Semrush tracked more than 230,000 prompts weekly from 14 July to 12 October 2025 across ChatGPT search, AI Mode and Perplexity. ChatGPT cited Reddit in close to 60% of responses in early August and around 10% by mid-September. Wikipedia fell from roughly 55% to under 20% over the same period, while both stayed steady on the other two engines.
The lesson is about stability, not about Reddit. A tactic built on one engine's current habits can stop working in a month, for reasons no outsider can see.
Widely repeated claims the evidence does not support
"Add an llms.txt file and AI tools will cite you." Google's guide says Search ignores the file and that it will "neither harm nor help". Ahrefs' log study of 137,210 domains found that 97% of published files were not requested once in May 2026, and that search-retrieval bots made up 1.1% of the requests that did occur. Details are in what llms.txt does and does not do.
"Schema markup makes AI cite your page." Ahrefs found that cited pages are far more likely to carry schema, then tested whether adding it changes anything. Across 1,885 pages and 4,000 matched controls, citations moved by +2.4% on AI Mode and +2.2% on ChatGPT, both indistinguishable from zero, and by -4.6% on AI Overviews. The study covered pages that were already heavily cited, so it does not rule out an effect on pages that are not yet seen at all. Google says structured data is not required for its AI features. Microsoft says schema helps its systems understand content. Keep schema for rich results. Do not buy it as an AI citation tactic.
"Content must be split into small chunks for AI." Google lists this as something site owners can ignore. Microsoft recommends clear headings, lists and tables. These are compatible: organise pages for readers. No independent test shows that chunking beyond that changes citations.
"More pages means more AI visibility." Page count had the weakest correlation of any factor in the Ahrefs brand study (0.17), and Google's guide says a high quantity of pages does not make a site more relevant.
"You must rank first on Google to be cited." The overlap figures above contradict this for every engine measured.
"We can track your position in ChatGPT." In the SparkToro and Gumshoe experiment, 600 volunteers ran 12 prompts 2,961 times across ChatGPT, Claude and Google's AI features. The chance of getting the same list of brands twice in 100 runs was under 1 in 100. The same data showed something more useful: how often a brand appears across many runs is fairly stable. One agency appeared in 85 of 95 responses to a single prompt. Appearance rate over repeated runs is a defensible measure. A position from a single run is not.
"This method gets you cited." No vendor offers that, and no study shows it. Google states that meeting every requirement does not mean a page will be crawled, indexed or served.
What a small business can reasonably do
In order of how well each step is supported:
Confirm that the search crawlers can reach your pages and that the content is in the server's HTML. This is the only step every vendor documents. An AI visibility audit starts here.
Check that the site is included in Google's Search generative AI control and is indexed with snippets allowed.
Publish pages that state specific, dated, sourced facts a model can lift. The controlled evidence, with its limits, points this way, and so does Google's advice on non-commodity content. This is the work described under answer engine optimisation .
Earn real mentions: directory listings, reviews, partner pages, press. It is the strongest correlate and the slowest to build.
Measure with the tools the engines provide and with repeated prompt runs, not single screenshots. See how to measure AI referral traffic in GA4.
If the immediate problem is that an assistant does not name your business, the checks are in why ChatGPT does not mention your business. WavX's wider offer is described under generative engine optimisation .
When this is not worth paying for
If your site has little search visibility and few mentions elsewhere, the evidence says the gap is in ordinary SEO and reputation, not in an AI-specific tactic. Google's guide says as much: optimising for its AI features is still SEO. A business that gets most of its work through referrals or marketplaces may see no return from any of this. And a tracking subscription makes little sense before the basic crawler and indexing checks have been done, which cost nothing.
Limits of this review
It covers published sources read on 2 October 2026. Engines change faster than studies are published. Two of the findings above reversed or shifted within a year.
Most independent data comes from companies with products to sell in this area.
None of the studies was designed around Indian queries or Indian businesses.
Several findings concern Google AI Overviews only and should not be generalised to ChatGPT or Perplexity.
No study measures enquiries or revenue. They measure citations and mentions.
Sources
Google Search Central: AI features and your website , read 2 October 2026
Google Search Central: Optimizing your website for generative AI features on Google Search , read 2 October 2026
Google Search Central: Spam policies for Google web search , read 2 October 2026
Google Search Console Help: Search generative AI control , read 2 October 2026
OpenAI: Overview of OpenAI crawlers , read 2 October 2026
Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? , read 2 October 2026
Perplexity: Perplexity crawlers , read 2 October 2026
Microsoft Advertising: Optimizing your content for inclusion in AI search answers (8 October 2025) , read 2 October 2026
Ahrefs: An analysis of AI Overview brand visibility factors, 75K brands (26 May 2025) , read 2 October 2026
Ahrefs: 76% of AI Overview citations pull from the top 10 (21 July 2025) , read 2 October 2026
Ahrefs: Update, 38% of AI Overview citations pull from the top 10 (2 March 2026) , read 2 October 2026
Ahrefs: Only 12% of AI cited URLs rank in Google's top 10 for the original prompt (11 August 2025) , read 2 October 2026
Ahrefs: Do self-promotional "best" lists boost ChatGPT visibility? (4 December 2025) , read 2 October 2026
Ahrefs: We tracked 1,885 pages adding schema. AI citations barely moved (11 May 2026) , read 2 October 2026
Ahrefs: We analyzed 137K sites: 97% of llms.txt files never get read (15 June 2026) , read 2 October 2026
Aggarwal and others: GEO: Generative Engine Optimization, KDD 2024 (arXiv 2311.09735) , read 2 October 2026
Semrush: The most-cited domains in AI, a 3-month study (10 November 2025) , read 2 October 2026
Profound: AI platform citation patterns (5 June 2025, updated August 2025) , read 2 October 2026
SparkToro: AIs are highly inconsistent when recommending brands or products (27 January 2026) , read 2 October 2026 from the Internet Archive copy dated 1 October 2026
Frequently asked questions
How does ChatGPT choose which websites to cite? OpenAI does not publish its selection rules. What it documents is the precondition: a page must be reachable by OAI-SearchBot to be shown in ChatGPT search answers. Published studies add that ChatGPT's cited pages overlap little with Google's top ten for the same prompt and that its mix of sources has shifted sharply within weeks.
Does schema markup help a page get cited by AI? The best available test says not measurably. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched control pages in a study published on 11 May 2026 and found no meaningful rise in citations on Google AI Overviews, AI Mode or ChatGPT. Schema is still worth keeping for rich results in ordinary search.
Do I need to rank in Google's top 10 to appear in AI answers? No, although ranking helps. In Ahrefs' March 2026 data, about 38% of pages cited in AI Overviews also ranked in the top 10 for the same query. For ChatGPT, Gemini and Copilot, an August 2025 study found only about 12% of cited links in Google's top 10 for the same prompt.
Can anyone promise that an AI tool will cite my business? No. Google states that indexing and serving are not assured even for pages that meet every requirement, Microsoft says no practice ensures selection, and a January 2026 experiment found that the same prompt almost never returns the same list of brands twice. Anyone promising a citation is promising something the engines do not offer.
What is the single best-supported thing to do? Make sure the search crawlers of the engines you care about can fetch your pages, because every vendor documents that as a requirement. After that, the factor with the strongest published correlation is being mentioned by name on other websites, which is slow and cannot be faked safely.
About the author
WavX Editorial Team
Engineering & delivery team, WavX Solutions
Written and fact-checked by the WavX Solutions engineering team in Gurgaon, Delhi NCR — the people who scope, price and ship these builds. Costs and timelines quoted here come from projects we have actually delivered, not vendor price lists.
All articles by WavX Editorial Team →
Build your own software — your way, your pricing.
WavX Solutions is here to create your own software in a fully custom way, built exactly how you work — with a pricing model that fits your business. Connect now and let's build it.
Contact Now helpwavx@gmail.com
More on AI Development
AI Development hub
AI Development AI Customer Support Agent Cost in India 2026: ₹3L–₹12L Real Pricing
AI Development How to Build an AI Agent for Your Business 2026: ₹5L–₹45L+ Development Guide
AI Development AI Tools Every Small Business Should Use in 2026
AI Development How to Add an AI Chatbot to Your Website (2026 Guide)
AI Development How to Automate Invoicing With AI (2026 Guide)
AI Development How to Build a Custom AI Agent for Your Business 2026
Keep reading
SEO vs GEO: How Search Is Changing & What to Do in 2026
Read Generative Engine Optimization: Get Cited by ChatGPT & Gemini
Read How to Send Automated WhatsApp Messages to Customers 2026
Read How to Set Up Email Automation for Your Business 2026
Read How to Use AI for Customer Support (2026 Owner Guide)
Read Generative Engine Optimization (GEO): Complete 2026 Guide
Read AI Automation for Business: Cost, ROI & Use Cases (2026)
Read AI Voice Agent & Calling Bot Development Cost in India (2026)
Read WhatsApp AI Chatbot Development Cost in India (2026)
Read RAG Chatbot Development Cost, Architecture & Guide (2026)
Read AI Agent Development Cost in India (2026): Complete Pricing Guide
Read How to Rank Higher on Google in 2026: SEO Basics for Businesses
Read