How to Get Cited by ChatGPT and AI Overviews (What the Data Actually Shows) | Simple AI Tools

How to Get Cited by ChatGPT and AI Overviews (What the Data Actually Shows)

88% of AI Citations Don't Come From Page One — Here's What Actually Gets You Cited

What published research shows about how ChatGPT and Google AI Overviews really choose their sources — and which popular tactic the data does not support.


Here is the finding that should reorganise how you think about this entire problem.

Ahrefs analysed 15,000 queries and compared the URLs that AI tools cite against Google's top ten organic results. The overlap was about twelve percent. Which means roughly 88% of AI citations point to pages that do not rank on page one for the query being answered.

Ranking well and getting cited are not the same achievement. They are only loosely correlated. And that single fact invalidates the most common assumption in this field — that AI visibility is a byproduct of good SEO, and that if you simply keep doing what you were doing, the citations will follow.

They may not. So this article works from published research rather than assertion: what the data shows about how these systems select sources, what separates a page that gets cited from a page that actually shapes the answer, and which widely-recommended tactic the evidence does not support.

Selection vs Absorption: The Two-Stage Model

A recent academic measurement study examined 602 controlled prompts across ChatGPT, Google's AI surfaces and Perplexity, covering more than 21,000 search-layer citations and over 18,000 successfully fetched pages. Its central contribution is a distinction most practitioners collapse into one thing.

Getting cited happens in two separate stages, and they have different requirements:

Stage one — citation selection

The platform decides to search, and then decides which sources to pull. Eligibility here rests on things like authority, recognisability, language, and whether your page fits the domain context of the query. This is the stage most SEO advice targets.

Stage two — citation absorption

Being selected is not the same as mattering. Absorption is whether your page actually contributes language, evidence, structure or factual support to the generated answer. A page can appear in the citation list and contribute essentially nothing.

This distinction has real consequences. The study found a sharp divergence between citation breadth and citation depth: Perplexity cites the most sources per prompt and Google also cites broadly, while ChatGPT cites fewer sources but shows substantially higher average citation influence among the pages it fetches.

Translated into strategy: on Perplexity, being present is a reasonable goal. On ChatGPT, presence without influence is close to worthless, because it cites less and leans harder on what it does cite.

What High-Influence Pages Actually Contain

The same study characterised the pages that scored highest on absorption. They shared four properties:

  • Longer. Thin pages get selected occasionally and absorbed rarely.
  • More modular. Content organised into discrete, self-contained units rather than continuous undifferentiated prose.
  • More semantically aligned with the answer being generated — the page is genuinely about the thing being asked, not adjacent to it.
  • Containing extractable evidence. Specifically: definitions, numerical facts, comparisons, and procedural steps.

That last item is the most actionable thing in this article. The researchers frame the whole discipline as evidence-container design: a page must first be eligible for selection through authority and context, then be useful — meaning it must contain discrete, liftable units of fact that a model can pick up and place into an answer.

A well-written page with no extractable evidence loses to a plainer page that states a definition cleanly, gives a number with context, or lays out a comparison in a table.

⚠️ A finding that contradicts common advice: the study reports that Q&A formatting alone does not improve absorption. Adding an FAQ block to a page that lacks substantive extractable evidence does not make it more likely to shape an AI answer. FAQ sections still serve other purposes — featured snippets, structured data eligibility, genuine reader utility — but the widespread belief that question-and-answer formatting is itself a GEO lever is not supported here.

⚡ Now for the Tactic Everyone Is Selling You

There is one recommendation that appears in almost every guide on this subject, sold by agencies as essential infrastructure and auto-generated by major SEO plugins.

Two independent traffic studies and Google's own documentation say it does approximately nothing for AI citations.

The numbers are below.

The llms.txt Myth, in Numbers

The proposal is reasonable on its face. An llms.txt file is a Markdown document at your site root offering models a clean, structured map of your important content, sparing them from parsing navigation, cookie banners and script bundles. It was proposed by Jeremy Howard in September 2024 to solve a real technical problem.

The question is whether anything actually reads it. Three independent lines of evidence say largely no.

Evidence one: Google says it does not use it

Google's guide on optimising for generative AI features, updated in May 2026, tells site owners directly that llms.txt is not needed for AI Overviews, AI Mode, or any other generative AI Search feature. It groups the file alongside content chunking, AI-specific rewriting and special schema as tactics that do not help. Earlier, Gary Illyes confirmed Google does not support it and has no plans to, and John Mueller compared it to the long-discredited keywords meta tag.

Evidence two: almost nobody requests the file

Ahrefs examined 137,000 domains in its Web Analytics data. Among sites that had a valid llms.txt file, 97% received zero requests for it during May 2026 — not from AI bots, not from anyone.

Evidence three: the crawlers that matter aren't fetching it

A separate analysis of more than 515 million LLM bot traffic events, filtered specifically to the user agents that drive AI citations — GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Google-Extended — found requests to /llms.txt to be a statistically negligible share of total traffic.

Where it genuinely does work

This is not a case of the file being useless. It is a case of it being useful for something other than what it is being sold for. AI coding agents — Cursor, Windsurf, Claude Code, GitHub Copilot, Cline, Aider — routinely fetch llms.txt when pointed at documentation sites. Anthropic recommends it in its guidance for agent-facing content, and OpenAI uses it within its own developer tooling.

The honest rule: if you run a developer documentation site or an API reference, build one. It is cheap and it has a real audience. If you run a blog, a local business or an ecommerce store hoping to appear in AI Overviews, this file is not your lever, and the hours are better spent elsewhere.


What Google Actually Tells Site Owners

Google's published guidance on generative AI visibility is notably unglamorous. It emphasises core fundamentals — original content grounded in real experience, sound technical health, genuine authority — and explicitly declines to endorse special markup, AI-specific rewriting, or content chunking as levers.

This is easy to dismiss as Google being cagey. But it aligns with the absorption research: what makes a page useful to a generative system is substantive, well-organised, evidence-dense content. There is no formatting trick that substitutes for having something extractable to say.

Why One Strategy Cannot Cover All Platforms

The platforms do not overlap nearly as much as people assume. Research into the state of AI citations found that only around 11% of domains are cited by both ChatGPT and Perplexity. Combined with the twelve percent Google overlap figure, the picture is of largely separate ecosystems rather than one AI search layer.

Platform Citation Behaviour What to Optimise For
ChatGPT Cites fewer sources, but each carries substantially more influence on the answer Absorption. Depth, evidence density, semantic alignment
Perplexity Cites the most sources per prompt Selection. Breadth of presence and eligibility
Google AI Overviews Cites broadly; only ~12% overlap with its own top-10 organic Fundamentals. Authority, technical health, original substance

One further pattern worth knowing: analysis of ChatGPT's citations by top-level domain shows commercial .com domains taking over 80% of citations, with .org sites second at roughly 11%. The often-repeated claim that AI systems overwhelmingly prefer institutional and non-profit sources does not hold up as a general rule — though it does hold within specific verticals such as healthcare, where institutional medical domains dominate heavily.

The Surface You Don't Control

A significant share of AI citations point to places you cannot publish on directly — community discussions, forums, and user-generated threads.

The detail that matters: analysis of ChatGPT's Reddit citations found that roughly 99% point to individual threads containing substantive discussion, rather than to subreddit landing pages, user profiles, or brand-authored posts.

This has an obvious wrong reading and a correct one. The wrong reading is to start seeding promotional threads, which is both against platform rules and empirically ineffective given that brand-authored content is not what gets cited. The correct reading is that genuine, substantive participation in real discussions — answering questions properly, in public, where your expertise is relevant — creates citable artefacts that your own domain cannot produce.

The Evidence-Based Checklist

Everything below traces to a finding in this article. Nothing is included because it is popular.

Do This Because
Include definitions, numbers, comparisons and procedural steps These are the extractable evidence genres associated with high absorption
Write modular, self-contained sections Modularity correlates with citation influence
Go deeper on fewer topics Length and semantic alignment both correlate with absorption
Track citations per platform separately Domain overlap between platforms is roughly 11%
Participate substantively in real communities Thread-level discussion is what gets cited, not brand posts
Skip llms.txt unless you publish developer docs 97% of files get zero requests; Google confirms it doesn't use it
Don't treat page-one rankings as an AI strategy 88% of AI citations come from outside the top ten

Frequently Asked Questions

Does ranking in Google get me cited in AI Overviews?

Only weakly. Ahrefs found roughly 12% overlap between URLs cited by AI tools and Google's top-ten organic results across 15,000 queries, meaning the large majority of AI citations go to pages that are not ranking on page one for that query.

Do I need an llms.txt file to get cited?

No, unless you publish developer documentation. Google's May 2026 guidance states the file is not needed for AI Overviews or AI Mode, an Ahrefs study of 137,000 domains found 97% of valid llms.txt files received zero requests in a month, and analysis of over 515 million bot events found citation-driving crawlers request it negligibly. It does have real value for AI coding agents reading documentation sites.

Does adding an FAQ section improve my chances of being cited?

Not by itself. Research into citation absorption found that Q&A formatting alone does not improve whether a page influences a generated answer. FAQ sections remain useful for featured snippets, structured data and reader experience, but formatting is not a substitute for substantive extractable content.

What is the difference between citation selection and citation absorption?

Selection is whether a platform chooses your page as a source. Absorption is whether that page actually contributes language, evidence or structure to the final answer. A page can be cited without meaningfully shaping the response, and the two stages have different requirements.

Should I optimise differently for ChatGPT than for Perplexity?

Yes. ChatGPT cites fewer sources but relies on them more heavily, which rewards depth and evidence density. Perplexity cites far more broadly, which rewards presence and eligibility. Only around 11% of domains are cited by both, so a single unified approach leaves most of the surface uncovered.

What kind of content is most likely to be absorbed into an answer?

Pages that are longer, modular, closely aligned with the query topic, and that contain extractable evidence — clear definitions, specific numbers, direct comparisons, and step-by-step procedures.

The Takeaway

Most advice on getting cited by AI systems is a formatting trick sold with confidence and no measurement behind it. The published research points somewhere less exciting and considerably more durable.

Being cited requires eligibility — authority, context, technical soundness. Being absorbed, which is what actually matters on the platforms that cite selectively, requires that your page contain discrete, liftable evidence: things stated clearly enough that a model can pick them up and use them.

That is not a hack. It is the same thing that makes a page useful to a human reader, which is probably why it survives every algorithm change. Write the definition properly. Include the number. Build the comparison table. Show the steps.

Then track your citations per platform, because they are not one audience and never were.

Sources referenced: Ahrefs query-overlap and llms.txt adoption analyses; Google's guidance on optimising for generative AI features (updated May 2026); on-the-record statements from Google Search staff; an academic measurement study of citation selection and absorption across 602 prompts; Profound's platform citation pattern research. The academic study is a preprint — treat its findings as strong directional evidence rather than settled consensus.

🚀 Stay Connected With Simple AI Tools

Evidence-based AI, SEO and automation guides — no hype.

👇 💬 Drop your comment below and let us know your thoughts! ✨

The AI Explorer

Written by

The AI Explorer

Contributor at Simple AI Tools, covering AI tooling, applied machine learning and developer workflows. Every tool featured here is tested hands-on before it is written about.

  • Hands-on tested
  • Independent reviews
  • Updated

Comments

Share