TL;DR
Wikipedia accounts for 5% of all ChatGPT citations and appears in 18% of conversations with any source, yet the citation economy is highly unequal with a Gini coefficient of 0.8, meaning a tiny set of domains hoard most of the trust. Adding statistics to a page is the single strongest lever for earning an AI citation, boosting visibility-weighted word count by 41%, while citing external sources can lift a fifth-ranked page by 115% but actually hurts the top-ranked page by 30%.
The five most-cited domains overall are Reddit, YouTube, LinkedIn, Wikipedia, and Forbes, but the mix diverges sharply by engine—ChatGPT leans on Wikipedia and Forbes, Google AI surfaces favor Facebook and Yelp, and Perplexity relies on Reddit and G2 for B2B queries. Your bottom line: build a competitor source map by tagging each citation as a comparison article, forum thread, review listing, or vendor page, because the fix for losing to a G2 listing (get listed and get reviews) is completely different from losing to a Reddit thread (show up authentically, don’t astroturf).
What a competitor source map is
Direct answer: An AI search competitor source map is a structured record of which specific pages an AI answer engine (ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Gemini) cites when it names your competitors in response to prompts relevant to your category — plus what those pages are (a comparison article, a review platform listing, a forum thread, a vendor's own docs), so you can tell why the engine trusted them and what evidence you'd need to displace or join them.
It is not a keyword report. A keyword rank tells you where a page sits in a list of ten blue links. A source map tells you which specific document an AI model pulled a claim from, how often that document (or its domain) shows up across your prompt set, and what kind of source it is. That distinction matters because, as the research below shows, ranking well in traditional search and getting cited in AI answers are now only loosely related.
Why citations have become the real competitive surface
Second, citation behavior is not evenly spread across the web — it clusters hard around a small set of domains and document types, and that cluster is different for every platform. Peec AI's analysis of 30 million cited sources across ChatGPT, Google AI Mode, Gemini, Perplexity, and AI Overviews (published March 31, 2026) found Reddit, YouTube, LinkedIn, Wikipedia, and Forbes as the five most-cited domains overall — but the mix diverges sharply by engine: ChatGPT leans toward Wikipedia and editorial sources like Forbes and TechRadar, Google's AI surfaces lean toward Facebook and Yelp, and Perplexity leans on Reddit and G2 for B2B queries (Peec AI, "Top domains cited by AI search: Analysis based on 30M sources," March 31, 2026).
Profound's separate analysis of roughly 730,000 U.S. ChatGPT conversations from October-December 2025 quantifies the concentration inside a single platform: Wikipedia alone accounts for 5% of all citations and appears in 18% of conversations that contain any citation, while the citation economy overall shows a Gini coefficient of 0.8 — high inequality, meaning a small number of domains absorb a disproportionate share of citations even though the long tail is wide (Profound, "How ChatGPT sources the web"). Profound's companion cross-platform study (August 2024-June 2025, 680 million citations) shows just how platform-specific the concentration is: Wikipedia makes up 47.9% of ChatGPT's top 10 sources, while Reddit makes up 46.7% of Perplexity's top 10 and a more distributed mix of Reddit (21.0%) and YouTube (18.8%) shows up in AI Overviews (Profound, "AI Platform Citation Patterns," 2025).
None of this is about keyword density. It's about which documents an engine has learned to trust for a given query shape, and that trust is measurable, engine-specific, and — as the next section shows — not permanent.
What actually earns a citation, per the research
Direct answer: The most direct evidence on why a page gets cited (rather than just which domains dominate) comes from the Princeton/IIT Delhi/Georgia Tech/Allen Institute "GEO: Generative Engine Optimization" paper, published at KDD 2024. The researchers built a benchmark of roughly 10,000 queries across nine datasets and tested nine content-level interventions against a generative answer pipeline. Two findings are directly useful for competitive positioning work:
- Adding statistics was the single strongest lever tested, improving a page's visibility-weighted word count by 41% and its "subjective impression" score by 37% (Aggarwal et al., "GEO: Generative Engine Optimization," 2024).
- Citing external sources ("Cite Sources") was weak alone but a strong multiplier in combination with other techniques (31.4% average lift) — and its effect depended heavily on where a page started: a fifth-ranked page gained 115.1% in AI visibility from adding citations, while the top-ranked page in the same query lost 30.3% of its share. Lower-authority pages have more to gain from evidence-led content than pages that already win on domain authority alone.
This is the mechanism behind a source map: if your competitor's cited page is a G2 comparison, a Reddit thread, or a specific data-backed report rather than their own homepage, the map tells you which evidence type is winning in your category — and the GEO paper tells you which interventions (statistics, quotes, sourcing) actually move an engine's trust in that direction.
Source types, by typical signal strength
| Source type | Where it dominates | Why engines trust it | How hard to earn |
|---|---|---|---|
| Reference/encyclopedic (Wikipedia) | ChatGPT, Perplexity | Structured, neutral, consensus-verified | Not directly earnable by brands |
| UGC/community (Reddit, forums) | Perplexity, AI Overviews | First-person, unpaid experience signal | Earned through genuine participation, not posting |
| Video (YouTube) | AI Overviews, AI Mode | Transcripts + engagement metadata | Requires real video content, not repurposed text |
| Review/comparison platforms (G2, Capterra, Yelp) | Perplexity (B2B), AI Overviews | Structured side-by-side data | Earned via listings, reviews, category presence |
| Editorial/news (Forbes, TechRadar) | ChatGPT | Third-party validation, "best-of" framing | Earned via PR, coverage, contributed pieces |
| Your own site/docs | All platforms, low share | Direct but treated as self-interested | Fully controllable, but weighted down as biased |
The practical implication: a source map isn't complete until you've tagged each competitor citation by which row it falls into, because the fix for "we're losing to a G2 listing" (get listed, get reviews) is completely different from the fix for "we're losing to a Reddit thread" (show up authentically in the conversation, don't astroturf it).
How to build one — step by step
- Define the prompt set. Pull 30-100 real prompts that a buyer would plausibly type into an AI assistant in your category — not head keywords, but full questions ("best X for Y," "X vs Z," "is X worth it"). Recommendation-intent phrasing ("best," "top," "alternatives") triggers citations far more reliably than single-word queries.
- Run the same prompt set across multiple engines. Citation composition differs enough by platform (Wikipedia-heavy on ChatGPT, Reddit-heavy on Perplexity, video-heavy on AI Overviews) that a map built from one engine will misdiagnose the problem.
- Capture the actual cited URL, not just the domain. A citation to a specific TechRadar comparison article behaves differently than a citation to TechRadar's homepage. Bare-domain tracking hides which page is doing the work.
- Classify each cited source by type using a scheme like the table above (reference, UGC, video, review platform, editorial, competitor-owned, your-owned).
- Tally frequency by competitor and by source type, not just by domain. You want to know both "who beats us" and "what kind of evidence beats us."
- Read the cited page itself, not just its metadata. Note what specific claim, statistic, quote, or structure the AI answer pulled from it — this is where the GEO paper's findings apply directly: is it winning because it has hard numbers, direct quotes, or structured comparison data?
- Sort gaps into "earnable" and "not earnable." You cannot become Wikipedia. You can become a page a review platform, journalist, or community cites. Spend effort where you can plausibly enter the source pool.
- Build evidence-backed content or listings against the specific gap, not a generic new blog post. If competitors win via G2 comparisons, the fix is a stronger G2 presence plus a comparison page with real numbers — not a keyword-optimized rewrite of your existing page.
- Re-run the prompt set on a fixed cadence. Citation composition moves — sometimes for content reasons, sometimes for reasons that have nothing to do with content quality (see below) — so a one-time map goes stale.
What this doesn't guarantee
Citation share is genuinely volatile, and not always for content reasons. Reddit's citation share on platforms like ChatGPT and Perplexity has swung sharply within a matter of months for reasons that have nothing to do with content quality — unannounced retrieval-mix changes on the model side, and legal disputes over data licensing and scraping on the business side. No source map can predict a lawsuit or an unannounced retrieval-mix change.
There's also a structural ceiling on diversity itself. A 2026 MIT study that ran 24,000 identical queries across 243 countries found that AI search "surfaces significantly fewer long tail information sources, lower response variety," and "significantly fewer high credibility information sources and significantly more low credibility information sources" than traditional search, with citations concentrating heavily toward a small set of top-1,000 sites (Aral, Li & Zuo, "The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale," MIT, Feb 13, 2026). That means for some queries there may simply be no realistic path to a citation regardless of how good your evidence is — the engine has already narrowed its trusted set before your content is even in consideration.
Finally, a source map describes correlation, not causation. Seeing that competitors win via G2 listings tells you G2 listings correlate with citations in your category on the engines you tested — it doesn't prove that adding a listing will itself produce a citation, since engine behavior, prompt phrasing, geography, and account state all shift outcomes independently of any single content change.
FAQ
Direct answer: No. Rank tracking tells you whether you appear; a source map tells you what's cited when a competitor appears, and what kind of evidence that citation represents. You need both, but they answer different questions — the source map is specifically diagnostic for "why do they win instead of us."
How many prompts do I need before the map is reliable?
Studies in this space use thousands to millions of prompts to establish stable domain-level patterns, but for a single competitive category, 30-100 realistic recommendation-intent prompts run across two or three engines is usually enough to spot the dominant source types. Treat single-prompt results as anecdote, not signal.
Does getting cited on Reddit or a review platform mean I should post there directly?
Not directly, and not synthetically. The research shows AI engines weight UGC because it reads as unpaid, first-person experience; a coordinated posting campaign is detectable and tends to be filtered or discounted, and platforms like Reddit have shown they can and will restrict access when they detect scraping or manipulation abuse. The earnable path is genuine participation and genuine reviews, not volume.
Why would a page ranking outside the top 100 in Google get cited in an AI Overview when my page ranking #3 doesn't?
Because AI Overviews increasingly pull from Google's internal "query fan-out" — related sub-queries the engine generates that you never typed — and a page that comprehensively covers one of those sub-angles can get cited even without ranking for your original term. This is the core finding behind Ahrefs' 2026 update showing top-10 overlap dropped from 76% to 38%.
Should I stop tracking traditional rankings and only watch AI citations?
No — the two are correlated but increasingly independent, and traditional rankings still feed AI Overviews' fan-out sourcing indirectly. Track both, but don't assume moving one moves the other.
How often should a competitor source map be refreshed?
Monthly at minimum for a fast-moving category, given the documented swings in platform citation behavior driven by unannounced retrieval changes and business disputes between engines and source platforms. A map older than a quarter should be treated as historical, not current.



