TL;DR

The most actionable number in all of Google's own AI Overviews documentation is that the Search Console generative AI report is the only first-party measurement tool worth building a workflow around — and it's rolling out gradually, not to every site at once. Independent research shows organic top-10 overlap with AI Overview citations varies wildly: Ahrefs measured it at 76% in mid-2025, then 37.9% six months later after a Gemini model update, while BrightEdge recorded both a 32.3% low and a 54.5% high across different verticals. A controlled difference-in-differences study found adding JSON-LD schema actually produced a 4.6% decline in AI Overview citations, not an uplift.

The 58% average CTR drop for the top organic result when an AI Overview is present (December 2023 to December 2025) is a more reliable number for content strategy than any citation-weighting theory. The article's verdict: treat any single claim about AI Overview citation patterns as provisional, run your own small-sample audits with matched controls, and accept that for most source-selection mechanics, the honest answer remains "we don't know."

Every few months a new post claims to have cracked the Google AI Overviews "algorithm" — a precise correlation coefficient here, a "78% more likely to be cited" stat there, usually with no visible sample, no dataset link, and no methodology section. Some of it is honest research. A lot of it is pattern-matching dressed up as a formula. The uncomfortable truth is that nobody outside Google's Search team knows exactly how AI Overviews select and weight sources, and the public studies that do exist frequently disagree with each other by wide margins.

That disagreement is itself useful information. This piece walks through what Google has actually documented, what independent research has measured (including where those measurements contradict each other), what you can responsibly conclude from your own sampling, and where the honest answer is still "we don't know."

What Google actually says

Direct answer: Google's own documentation on this is thinner than most SEO content implies, but it isn't silent. The core reference is AI Features and Your Website in Search Central, which states plainly: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." To be eligible as a supporting link, a page simply needs to be indexed and eligible to appear in classic Search with a snippet — there's no separate technical bar.

Google also describes the retrieval mechanism it uses to expand a query before generating a response, a technique it calls "query fan-out": while a response is being generated, "advanced models identify more supporting web pages, allowing us to display a wider and more diverse set of helpful links associated with the response than with a classic web search." Search Engine Land's guide to query fan-out is a useful plain-language expansion of this: Google issues multiple related sub-queries behind a single search — roughly the equivalent of you asking one question but Google's systems retrieving evidence for several adjacent questions at once, then synthesizing across all of them.

That single mechanic explains most of what looks confusing about AI Overview citations at first glance: a page can get cited for a query it doesn't rank for at all, because it was pulled in to answer a sub-query the searcher never typed. Google also rolled out a dedicated generative AI performance report in Search Console, giving site owners impressions and page-level visibility specifically for AI Overviews and AI Mode, separate from the standard Performance report — this is the one first-party measurement tool worth building a workflow around, and it's rolling out gradually rather than to every property at once.

Google's Guide to Optimizing for Generative AI Features reiterates that existing SEO fundamentals — crawlability, indexing, genuinely helpful content — remain the operative advice. It does not publish, and has never published, a ranking formula or citation-weighting model for AI Overviews specifically.

What independent research has found — and where it disagrees

Direct answer: Several research groups have run large-scale citation audits, and their headline numbers move around enough that treating any single figure as settled fact would be a mistake. The clearest example is the "how often does an AI Overview citation also rank in the top 10 organically" question:

StudySampleTop-10 overlap findingNotes
Ahrefs, mid-2025Earlier citation dataset~76%Cited widely as evidence AIO tracks organic rank closely
Ahrefs, Jan 2026863K keyword SERPs, 4M AIO URLs37.9% top 10; 31.2% positions 11–100; 31.0% beyond position 100Ahrefs attributes part of the drop to improved citation-parsing, part to Gemini 3 and heavier query fan-out — explicitly flags the two datasets as "not directly comparable"
BrightEdge, 16-month tracking (May 2024–Sep 2025)9 industriesOverlap grew from 32.3% to 54.5%Long-run trend toward convergence; healthcare/insurance/education 68–75%, e-commerce flat near 23%
BrightEdge, stricter top-10-only methodologySeparate pass~17%Same organization, different measurement window and inclusion rules than the 54.5% figure

See Ahrefs' write-up, the Search Engine Journal summary of the drop, and BrightEdge's 16-month overlap analysis.

The point isn't that one number is right and the others are wrong. It's that "top-10 organic overlap" isn't a stable, single-valued thing — it depends on the query set, the vertical, the citation-detection method, the time window, and which version of the underlying model Google is running that month. Anyone quoting one of these percentages as a universal constant is skipping the methodology section.

A few findings are more directly useful because they came from controlled before/after tests rather than raw citation counts. Ahrefs tracked 1,885 pages that added JSON-LD schema markup between August 2025 and March 2026 against 4,000 matched control pages, using a difference-in-differences design. Result: AI Overview citations for the schema-adding pages actually declined 4.6% (statistically significant but small), while AI Mode and ChatGPT showed changes indistinguishable from noise. Their conclusion: "adding schema produced no major uplift in citations on any platform." That's a useful corrective against the common claim that structured data is a citation lever — at minimum, it isn't a reliable, sizeable one on its own.

Separately, Ahrefs' CTR analysis of 300,000 keywords found a 58% average drop in click-through rate for the top-ranking organic result when an AI Overview is present, comparing December 2023 to December 2025 — a large, consistently replicated effect across several independent measurements cited in that piece, and arguably more load-bearing for content strategy than any citation-weighting theory, because it changes the economics of ranking #1 regardless of whether you're cited in the Overview itself.

A sampling method you can actually run

Direct answer: Given how much published research disagrees, the more defensible move for a specific business is to build your own small, honest dataset rather than importing someone else's percentage. A workable protocol:

  1. Fix a query set before you start. Pull 30–100 queries that matter to your business — a mix of head terms, long-tail how-to queries, and comparison queries. Lock the list before observing results, so you're not unconsciously cherry-picking queries that confirm a pattern.
  2. Capture consistently, not opportunistically. Log whether an AI Overview appears, which URLs are cited, their position in the citation list, and the organic ranking (if any) for the same URL on the same day. Repeat on a fixed cadence — weekly is reasonable — because AI Overview presence and citation sets are not stable; the same query can show a different Overview, or none at all, day to day.
  3. Record what you can't control for. Note query intent type (informational, comparison, transactional), whether the query triggered AI Mode's "Show more" expansion, and roughly how many citations appeared. These are the variables most likely to explain why your citation set doesn't match a published study's.
  4. Cross-reference with Search Console. Where the generative AI performance report is available on your property, compare its impressions/page data against your manual sample — it's Google's own count, not an inferred one from scraping.
  5. Separate correlation from cause explicitly. If cited pages in your sample tend to be longer, more structured, or more recently updated, write that down as an observed correlation in your dataset, not a mechanism. The schema study above is a good reminder that a plausible-sounding lever can turn out to do nothing, or the opposite of what's expected.

What's genuinely observable vs. genuinely unknown

Direct answer: Observable, with the caveats above: AI Overviews frequently cite pages outside the organic top 10, and the share doing so appears to be rising as query fan-out plays a larger role; certain domains (YouTube, Wikipedia, Reddit, and large publishers) recur disproportionately across many independent citation audits; AI Overview presence measurably suppresses top-result CTR; and a page does not need special markup or an "AI-specific" technical setup to be eligible — ordinary indexability is the bar Google states.

Genuinely unknown, and worth being blunt about with stakeholders: the actual internal weighting Google's ranking and retrieval systems apply to any single signal (freshness, structure, authority, brand mentions, entity coverage); how much of citation selection is query-specific fan-out versus static page quality; how personalization, geography, or account signals affect what an individual searcher sees; and whether findings from one vertical (say, comparison shopping) generalize to another (say, regulated financial content). Nobody publishing a "GEO ranking factors" listicle has access to Google's model internals, and treating vendor research as if it settles these questions — rather than as one data point among several disagreeing ones — is the exact overclaiming this piece is arguing against.

Turning findings into a content plan

Direct answer: None of this uncertainty means the exercise is pointless — it means the output should be framed as evidence-led hypotheses, not rules. A reasonable way to convert a sampling exercise into action:

  • Prioritize the sub-query surface, not just the head term. Since fan-out means Google may retrieve evidence for questions adjacent to what was typed, build content that explicitly and separately answers the two or three most obvious follow-up questions a searcher would have, rather than optimizing narrowly for one keyword.
  • Treat "not in the top 10" as normal, not disqualifying. Given that roughly a third to two-thirds of citations in various studies come from outside the top 10, don't assume a page needs to win the organic race before it can be cited — instead track whether it's getting cited at all, using it as its own signal.
  • Retest before committing budget to a specific tactic. If a competitor claims schema markup or a particular structure boosted their citations, treat that as a hypothesis to validate on your own query set before rolling it out site-wide — the schema study above is a direct example of a widely repeated claim that didn't hold up under a controlled test.
  • Watch CTR economics alongside citation counts. A page can be cited in an AI Overview and still see organic traffic fall, because the Overview itself is absorbing clicks. Report both numbers together so a "win" in citation share isn't mistaken for a traffic win.
  • Re-run the sample on a cadence. Because citation sets are volatile and Google's underlying models change (Ahrefs explicitly attributes part of its 76%→38% swing to a model version change), a one-time audit goes stale within a quarter. Build this as a recurring check, not a report you produce once.

The honest version of this work looks less like "we reverse-engineered the algorithm" and more like a lab notebook: here's what we sampled, here's what we saw, here's what changed since last time, and here's what we still can't explain. That's a less exciting pitch than a secret formula, but it's the version that won't need walking back next quarter.

Sources: