Two pages can cover the same topic, run the same length, and sit next to each other in Google, and an AI answer will still cite only one of them. The deciding factor is rarely the subject. It is how easily a model can reach the page, lift a passage that stands on its own, and find corroboration for what that passage claims.
Unusually for this field, there is published research to work from rather than agency folklore: a peer-reviewed paper from Princeton, large-scale citation studies from Ahrefs and Semrush, and Google’s own documentation on what makes a page eligible. This piece breaks the selection process into five signals, each with the evidence behind it, and then looks at the awkward part — most AI citations go to a handful of sites that are not yours.
How does AI choose which sources to cite?
AI systems filter sources on five things at once: whether a crawler can fetch the page, whether it answers the exact question asked, whether a self-contained passage can be lifted from it, whether its facts agree across the web, and whether anyone else vouches for the brand. Pass four and fail one, and you are usually out.
None of the five works in isolation. Excellent content on a page blocked to AI crawlers is invisible; a perfectly crawlable page of marketing vagueness has nothing worth quoting. The order below is roughly the order in which a page gets eliminated.
Signal 1: crawlability — can the model reach the page at all?
This is a gate, not a ranking factor. Google states plainly that “to be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet,” per its AI features documentation. Fail that, and nothing else you do counts.
Two things commonly break this. The first is robots.txt. In its July 2025 analysis, Cloudflare found that only about 37% of the top 10,000 domains even have a robots.txt file, and that GPTBot is disallowed in 7.8% of the ones that do — often by a plugin default or a hosting setting nobody chose deliberately. The second is rendering: if your prices and service descriptions only appear after JavaScript runs, a fetcher reading raw HTML may see an empty shell.
Worth noting what is not required. Google is explicit that “there are no additional requirements to appear in AI Overviews or AI Mode” and that “you don’t need to create new machine readable files, AI text files, or markup” for these features. If you have been told an llms.txt file is the fix, we looked at what that file actually does in what is llms.txt and do you need it.
Signal 2: relevance — does the page answer the exact question asked?
The question a model is answering is often not the question the user typed. Google describes a “query fan-out” technique, where AI Overviews and AI Mode issue multiple related searches across subtopics and then pull supporting pages from those. Your page competes for the sub-question, not the headline query.
The consequence shows up in the data. Ahrefs analysed 863,000 keyword SERPs and 4 million AI Overview URLs and found that 37.9% of cited URLs appeared within Google’s top 10 result blocks (Linehan and Guan, March 2026) — down from roughly 76% in their July 2025 study. Ahrefs attributes the shift to fan-out. Read it either way and the conclusion holds: a majority of AI Overview citations now come from pages that were not ranking on page one for the original query.
That is genuinely good news if you are not a market leader. It means the citation is going to whoever answered the narrow sub-question best, not automatically to whoever ranks first. It also means a page that buries its answer under three paragraphs of positioning has nothing to match against.
Signal 3: quotable passages — can a model lift a self-contained line?
This is the one signal with a controlled experiment behind it. In GEO: Generative Engine Optimization (Aggarwal et al., presented at KDD 2024), Princeton researchers tested nine content modifications across a 10,000-query benchmark and measured how visible the modified source became in the generated answer.
Three edits won: adding relevant statistics, adding credible quotations, and citing reliable sources. The paper reports these delivered “a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric.” Two things did not work. Keyword stuffing — the paper’s words — offered “little to no improvement.” Nor did simply adopting a more authoritative, persuasive tone: the researchers found “no significant improvement, demonstrating that Generative Engines are already somewhat robust to such changes.”
The most useful table in that paper is the one nobody quotes. Broken down by the source’s original search ranking, the gains landed almost entirely on the underdogs: the fifth-ranked source gained 115.1% visibility from citing sources and 97.9% from adding statistics, while the top-ranked source lost ground on the same edits (−30.3% and −20.6%). The authors put it plainly: “websites that are ranked lower in SERP, which typically struggle to gain visibility, benefit significantly more from GEO than those ranked higher.”
One honest caveat, because the number gets passed around without it: this study ran in 2023 on a generative engine built from GPT-3.5-turbo over the top five Google results. Today’s systems are different. Treat the direction as well-evidenced and the exact percentages as historical.
The practical test takes ten seconds. Cut a sentence out of your page and read it alone. “Pricing depends on a number of factors” tells a model nothing. “A single-location dental practice site costs PLN 6,000–9,000 net and takes four weeks” survives extraction intact.
Signal 4: consistency — do your facts agree with themselves?
Models cross-check. When your address, hours, service list, or prices differ between your site, your Google profile, and a directory, that inconsistency is a reason to cite the competitor whose facts line up. It is not a small nicety; it is the difference between a verifiable claim and an unverifiable one.
The same applies inside a single page. Google’s structured data policies require that “your structured data must be a true representation of the page content” and instruct you not to “mark up content that is not visible to readers of the page.” Schema saying you are open until 6pm while the page says 5pm is not a rounding error to a system whose job is deciding what it can safely repeat.
Signal 5: authority — who else on the web vouches for you?
Here is where the English-language picture differs most from the tidy version of this advice. Ahrefs studied 75,000 brands and correlated search metrics against brand mentions in ChatGPT, AI Mode, and AI Overviews (Linehan and Guan, December 2025). The strongest correlations were all off-site: YouTube mentions at 0.737, branded web mentions at 0.656–0.709 depending on platform, and branded anchors at 0.511–0.628. Backlinks sat around 0.20–0.30.
Read that carefully. Mentions of your brand elsewhere on the web correlated roughly two to three times more strongly with AI visibility than link counts did. The authors are careful, and so should we be: they state that “correlation isn’t causation” and that improving these metrics will not automatically boost AI visibility. But the pattern points somewhere specific — the question is less “is my page readable” and more “where does independent evidence of my existence live.”
This is the slowest of the five signals to move, which is exactly why it should start in parallel with the others rather than after them.
Where do AI citations actually come from?
Mostly not from sites like yours. Peec AI analysed 30 million cited sources and found Reddit the most-cited domain across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, followed by YouTube, LinkedIn, Wikipedia and Forbes (reported by Search Engine Land, March 2026). Community platforms and reference sites take a large share of the pool.
That reframes the goal. If half the citation slots go to Reddit threads, LinkedIn posts and YouTube transcripts, then being mentioned on those surfaces is part of being cited — not a separate social media task. It also explains why the Ahrefs correlations point off-site.
The per-platform picture is worth knowing because it is unstable. Semrush tracked 230,000 prompts and over 100 million citations in weekly snapshots between 14 July and 12 October 2025 across ChatGPT search, Google AI Mode and Perplexity:
| Aspect | Google AI Mode / AI Overviews | ChatGPT search | Perplexity |
|---|---|---|---|
| Source pool | Google’s index; page must be indexed and snippet-eligible (Google) | live web search plus training data | its own crawler and index |
| Observed favourites | LinkedIn in nearly 15% of responses; Wikipedia in only ~2% (Semrush) | Wikipedia and Reddit dominant before September 2025 (Semrush) | Reddit, LinkedIn and NIH among top sources (Semrush) |
| Stability | AI Overview citations from the organic top 10 fell from ~76% to 37.9% between July 2025 and March 2026 (Ahrefs) | Reddit citations fell from ~60% of responses to ~10%, Wikipedia from ~55% to under 20%, within weeks in September 2025 (Semrush) | comparatively steady over the same window (Semrush) |
The takeaway is not to chase a platform’s current favourites. It is that any strategy built on one system’s habits has a short shelf life, which is why the five signals — the things all of them check — are the durable part. We cover platform-by-platform differences in AI visibility vs SEO and the vocabulary in GEO vs AEO vs SEO.
What to do with this on your own site
The five signals are also five workstreams, ordered here from cheapest to slowest:
- Crawlability — hours. Check
robots.txtfor GPTBot, ClaudeBot and PerplexityBot, and confirm key content exists in raw HTML. - Relevance — days. Give each real customer question its own heading, answered in the first two sentences.
- Quotability — days. Put numbers, prices, timelines and named sources into plain text that survives being cut out of context.
- Consistency — one cleanup, then discipline. Identical facts everywhere a machine reads them, schema matching visible content.
- Authority — months. Reviews, mentions, and presence on the platforms that already own the citation pool.
Before changing anything, take a baseline, or you will not be able to tell a real improvement from the day-to-day variance the Semrush data makes obvious. How to check if AI recommends your business covers the manual version; how to measure AI visibility covers reading a trend instead of a snapshot. The full implementation list lives in how to optimize your website for AI search, and if you want to work through it systematically, there is an AI SEO audit checklist.
If you would rather know which of the five signals is failing on your site and in what order to fix them, that is what our AI visibility audit does: 39 criteria across six areas, the same set of customer questions asked across ChatGPT, Gemini, Perplexity and Claude, and a scored PDF with at least five prioritised fixes — or a full refund. It starts at PLN 499 net (around $125). Prefer to talk first? Book a free 20-minute consultation and we will tell you honestly whether it applies to you, or start broader with the free 3-minute marketing audit.
FAQ
Does AI quote my page word for word, or paraphrase it? Usually it paraphrases, keeping the meaning and linking back to the source. Verbatim quoting is most likely for short, specific facts — a price, a date, a measurement — where rephrasing would change nothing. Either way, the passage has to be extractable, which is why self-contained sentences outperform hedged ones.
Does a longer page have a better chance of being cited? Length itself is not the variable; density of quotable statements is. The Princeton GEO study’s best-performing edits were adding statistics, quotations and citations — not adding words. A short page with three precise, sourced answers gives a model more to work with than a long one where the specifics are buried in generalities.
Does AI prefer newer content? It depends on the question. For anything that changes — prices, regulations, availability — freshness matters and systems try to account for it. For stable topics like definitions and mechanisms, precision beats publication date. What the citation studies do show is that the systems change fast, even when your content does not.
Should I be trying to get mentioned on Reddit? Not by posting promotional threads, which communities remove and which would fail the consistency test anyway. The finding to act on is that third-party validation carries weight: genuine reviews, being named in industry and local sources, and being discussed where your customers already are. Ahrefs’ 75,000-brand study found off-site brand mentions correlated far more strongly with AI visibility than backlinks did — though the authors note correlation is not causation.
Can I pay to be cited? No. There is no advertising slot inside a composed AI answer the way there is above Google’s results. Citations are earned through readability, consistency and third-party evidence — which is why smaller companies that do good work can compete here, once the work is packaged in a form a machine can verify. We unpack that gap in why ChatGPT doesn’t recommend your business.