A page you know that ranks in the top three for a query gets completely ignored by the AI Overview sitting above it, while a less authoritative article from a domain you’ve never heard of gets quoted verbatim. Why?
If you’ve been working in SEO the last handful of years like I have, you’ve watched it happen. It happens enough that it’s justifiable to call it a pattern. And the frustrating part isn’t that the pattern exists, it’s that it’s very hard to get a clear account of what drives it.
This article aims to be that account. Not a complete one, since Google hasn’t published a sourcing algorithm we can audit, but a grounded one, built from what the data shows about which content gets quoted, which gets buried, and what structural and editorial choices seem to move the needle either way.
Ranking first doesn’t necessarily mean getting quoted
The most disorienting thing about AI Overviews, for anyone who has spent years thinking in terms of organic rankings, is that the two systems (Overviews and search algorithms) run on genuinely different logic. As of April 2026, Search Engine Land reported that the overlap between AI Overview citations and organic rankings grew from 32.3% to 54.5% between May 2024 and September 2025.
That improvement sounds reassuring until you read it the other way: Even at peak convergence, nearly half of AI Overview citations came from pages that were not ranking in the top organic positions for the same query. Ranking well and getting cited are related, but they’re not the same thing.
Google’s own guidance offers a starting point but not an answer. Google Search Central’s documentation describes a process of “grounding” responses in web content, framing it as an extension of how Search has always evaluated quality. What it doesn’t tell you is why one well-ranked, well-structured page gets pulled into an Overview while an adjacent competitor doesn’t. The Google I/O 2024 keynote emphasized that AI Overviews are designed to synthesize information across multiple sources rather than simply promote the highest-ranked one. That’s the closest Google has come to officially acknowledging that citation and ranking are separate decisions.
This creates two failure modes that content publishers fall into with almost equal frequency.
- Assuming AI Overview inclusion tracks directly with organic rank; that if you fix your rankings, the citations follow automatically.
- Concluding that the whole thing is an arbitrary black box not worth trying to understand.
The first assumption leads you to ignore citation structure entirely, and the second leads you to stop trying to understand it at all. Semrush’s AI search analysis documented a 527% increase in AI search traffic in a single year, which means the stakes of getting this wrong are compounding fast. And BrightEdge data from early 2026 shows AI Overviews now trigger on nearly half of all tracked queries. AI Overviews are (becoming?) the default experience for the kinds of informational queries your content was built to answer.
The selection logic is neither random nor fully transparent, but it is not unknowable either. Enough patterns have emerged in practitioner research, from Ahrefs, Semrush, BrightEdge, and independent analysts, that we can make evidence-grounded inferences about what the selection logic rewards. Ahrefs found that AI Overviews reduced clicks to top-ranking content by 34.5% even as Google publicly stated the opposite.
If you’ve already absorbed the traffic losses from zero-click results (which it’s likely you have), AI Overview exclusion adds another layer: not just fewer clicks, but no attribution, no brand signal, no presence in the answer at all. You disappear from the conversation entirely.
Discrete, attributable facts get cited
If you take one thing away from this article, let it be this:
The clearest pattern in AI Overview citation behavior is that AI systems select passages that work as standalone answers.
A piece of content has high claim density when individual sentences or short paragraphs contain a complete, verifiable assertion that can be lifted without the surrounding context and still mean something. “Adults need 150 minutes of moderate aerobic activity per week” is a high-density claim: It names a quantity, a behaviour, and a timeframe. An AI system grounding a response about exercise recommendations can extract it, attribute it, and move on. A 1,200-word wellness article that builds toward the same point across four narrative sections can’t be extracted the same way, because no single sentence in that article carries the full informational weight.
The format implications are direct. Structured FAQs, definition-first explanations, numbered steps, and listicles all have naturally high claim density. Each item stands alone. Long introductions, contextual scene-setting, and paragraphs that build an argument across multiple sentences have low density, because meaning is distributed rather than concentrated.
Please don’t read this as me saying that listicles are better content – I’m just saying they’re more extractable content, which is different.
A short FAQ answer built around a single verifiable claim can get cited in an AI Overview while a long authoritative guide on the same topic gets passed over entirely. The guide may rank higher organically, and it almost certainly took more expertise to write, but if no individual sentence in it functions as a self-contained answer, AI will pass over it.
Users who encounter an AI summary are less likely to click through to source pages, which means the citation itself is often the only moment a publisher’s work is visible. If your content’s most citable sentence is buried in paragraph eleven, that moment never arrives.
There’s a real tension here that most content advice glosses over. High claim density content is often what traditional SEO dismisses as thin: short, direct, and declaratively structured. The content formats that score best on depth, dwell time, and topical authority metrics are exactly the formats that resist extraction. Our article on how LLMs process and cite web content points to the same principle, which is that legibility to a machine is not the same as depth for a human reader, and optimising for one without considering the other increasingly represents an opportunity for improvement.
Backlinks don’t buy you into AI overviews
The working assumption in SEO has always been that authority flows from links. High domain authority, strong backlink profiles, and years of accumulated PageRank are the signals that determine which pages win competitive queries. Link-based authority was always a proxy, but it was a consistent enough proxy that you could build a decade of strategy on it.
But there’s now data showing that traditional SEO metrics, including backlinks and domain authority, predict only 4% to 7% of AI citation behavior. The signals that have dominated search strategy for two decades are now nearly irrelevant to whether your content gets quoted in an AI Overview.
What appears to matter instead is source-type matching, the degree to which a piece of content resembles the kind of source a well-informed person would expect to answer that specific query:
- A “what is” query draws from encyclopedic or definitional sources.
- A “how to” query draws from instructional content with numbered steps and clear procedural structure.
- A “should I” query, in for example medical or financial contexts, draws from professional advisory sources with stated credentials and institutional backing.
Semrush’s analysis of the most-cited domains in AI systems found that citation patterns shifted across query categories, with Wikipedia and Reddit losing ground even as specialist sources held or gained position in their respective domains. The implication is that a local authority on a narrow topic, like a specialist clinic writing about a specific condition or a niche trade publication covering a specific industry process, can surface over a high-DA generalist publication that covers the same topic.
Domain authority measures accumulated reputation across the whole site, regardless of what the particular page is about. Source-type fit measures whether this page, on this topic, from this kind of publisher, is the right match for this query. AI Overviews appear to prefer the specific.
The authority signal that most consistently correlates with AI inclusion is not how many sites link to you, but how clearly your content declares who wrote it, what qualifies them, and what the content is about. Of these signals, the one that appears to matter most is the simplest: a byline linked to an author page with stated expertise, because that is the signal an LLM can easily parse and verify.
Google’s own documentation on how AI features evaluate source content points toward trustworthiness and authoritativeness as key considerations, and the practical expression of those qualities is structural, not just reputational. An anonymous post on a high-DA site gives an AI system less to work with than a bylined piece on a mid-tier site where the author’s credentials are explicitly stated and linked.
What gets deprioritised is equally instructive.
Content that looks promotional, even when it contains accurate claims, is a consistent liability. A statistic buried inside a product page, for example, resists citation because the surrounding intent is transactional rather than informational. Thin topical coverage is another signal the grounding logic appears to penalise, which is why a specialist site with deep, consistent coverage of a narrow domain tends to outperform. Content whose authority signals are structurally legible, with proper heading hierarchy, clear author metadata, and schema that reflects semantic compatibility standards, gives AI more confidence in attribution.
What AI parses before it quotes you
Heading hierarchy
A properly nested heading hierarchy tells a parsing system which claims belong to which topic. When an AI encounters a factual assertion, it doesn’t read the page the way a human does, with accumulated context and inference. It needs structural anchors. An H2 or H3 sitting immediately above a paragraph signals that the content below it is about a specific, bounded subject. That signal allows the system to extract a claim and attribute it correctly, to associate a specific statistic or definition with the right context rather than a vague section of undifferentiated prose.
Ahrefs analyzed 6 million URLs and found that schema markup is more common on pages cited by AI Overviews than on pages that are not, a pattern consistent with the broader principle that machine-readable structure predicts citation inclusion more reliably than content quality alone.
Semantic HTML
The structural argument extends well below the heading level. Content marked up with semantic HTML elements, like article, section, aside, figure with figcaption, or definition lists where terminology is being defined, gives a parsing system explicit signals about the role each piece of content plays. Prose sitting inside generic div elements carries no such signals. The content may be identical word for word, but one version tells the system “this is a self-contained article with a discrete topic and a figure that illustrates a specific claim,” while the other offers nothing but text. The same logic applies to FAQ schema, HowTo schema, and Article schema with author metadata: in addition to helping Google index your content, these label the content type in a way that allows AI to match the content to a query intent with greater confidence.
The pattern that appears most consistently in AI Overview-cited content is specific enough to act on: a keyword-aligned H2 or H3 immediately above the cited passage; a short paragraph, typically under 100 words, containing the core assertion without promotional framing; and structured data markup that explicitly identifies what kind of content the page contains. Publishers who hit all three of those conditions are building content that is genuinely easier to read, extract, and trust.
It turns out that WCAG-compliant content structure is not a separate discipline from AI legibility. It pays to make your content accessible.
Proper alt text, logical reading order, labelled form elements, descriptive link text, correct heading nesting… These requirements exist because people who navigate by keyboard or rely on screen readers need content that communicates its structure explicitly. AI extraction systems have the same dependency. A page that’s inaccessible to a screen reader is, by the same structural logic, harder for AI to parse with confidence. Publishers who have invested seriously in accessibility compliance have, often without realising it, built content architecture that AI systems can attribute and quote more reliably than content from competitors who never made that investment.
The practical implication for content audits is significant. It’s now possible to evaluate content for AI extractability using substantially the same criteria applied in an accessibility audit: heading structure, semantic HTML, content segmentation, metadata completeness, alt text coverage. That gives content teams a concrete diagnostic framework rather than a vague instruction to “create better content.”
What AI Overviews tell you about content legibility
Think of AI Overview exclusion as a diagnostic, not a verdict. If you take the ten pages you most want cited and run them through a structured audit, the pages that score poorly on those dimensions are probably not failing an AI test, but a clarity test that predates AI Overviews entirely.
If an AI system can’t extract a clear, attributable claim from a page, neither can a journalist writing a round-up, a colleague searching your own CMS, or a reader scanning for a quick answer. The AI Overview is simply the most visible and measurable place where that failure now shows up.
The forward implication is that, as AI agents become more prevalent in search, the content that gets surfaced and acted on will be content that’s structurally legible to machines. And said machines reward the same practices that have always made content worth reading: clear structure, attributed claims, honest framing, and accessible presentation.
Which brings us to an uncomfortable truth: the content that survives Google’s AI Overview filter is not always the best content on a topic. It’s often just the most legible content on that topic. Those two things should be the same, but making them the same is the actual optimization problem worth solving. And it starts with the page, not the algorithm.
References
- AI features and your website | google search central | documentation | google for Developers. (2025a). In Google for Developers. https://developers.google.com/search/docs/appearance/ai-features
- AI overviews at the One-Year mark: Presence, size, and what they’re citing. (2025b). In Brightedge.com. https://www.brightedge.com/resources/weekly-ai-search-insights/ai-overviews-one-year-presence-size-citing
- AI search in 2025: Three key insights from BrightEdge’s AI overview and ChatGPT analysis. (2025c). Brightedge.Com. https://www.brightedge.com/blog/ai-search-2025-three-key-insights-brightedges-ai-overview-and-chatgpt-analysis
- Chapekis, A., & Lieb, A. (2025). Google users are less likely to click on links when an AI summary appears in the results. In Pew Research Center. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
- Harsel, L., Aleksandr Drozdov, & Skopec, C. (2025). The Most-Cited domains in AI: A 3-month study. In Semrush Blog. Semrush. https://www.semrush.com/blog/most-cited-domains-ai/
- Healthcare and AI overviews: How google sharpened its approach over three years. (2023). In Brightedge.com. https://www.brightedge.com/resources/weekly-ai-search-insights/healthcare-ai-evolution-google-2023-2025
- Jeffrey, M. (2026, April 2). Why your content doesn’t appear in AI overviews (even if it ranks in the top 10). Search Engine Land. https://searchengineland.com/why-content-doesnt-appear-in-ai-overviews-473325
- Law, R. (2026). Update: AI overviews reduce clicks by 58%. In SEO Blog by Ahrefs. https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/
- Measuring Google AI overviews: Activation, source quality, claim fidelity, and publisher impact. (2026). Arxiv.Org. https://arxiv.org/html/2605.14021v1
- Paruch, Z., & Skopec, C. (2025). 26 AI SEO statistics for 2026 + Insights they reveal. In Semrush Blog. Semrush. https://www.semrush.com/blog/ai-seo-statistics/
- Ray, L. (2025, December 6). Tech SEO connect 2025: Summary & takeaways. Lily Ray. https://lilyray.nyc/tech-seo-connect-2025-summary-takeaways/
- Sundar Pichai. (2024). Google I/O 2024: An I/O for a new generation. In Google. https://blog.google/intl/en-africa/products/explore-get-answers/google-io-2024-an-io-for-a-new-generation/

