TL;DR
Adding quotations from authoritative sources and relevant statistics produced the largest visibility lifts in a 10,000-query benchmark, while keyword stuffing was the only tactic that measurably hurt (Aggarwal et al., arXiv). Across 75,000 brands, branded web mentions correlated with AI answer visibility roughly three times more strongly than backlink counts did (Ahrefs).
Third-party coverage—reviews, comparison sites, press—accounts for the vast majority of citations, not a brand's own homepage. The article's verdict: stop mapping source pools; instead, run a fixed, repeated set of buyer prompts against your brand and confirmed competitors, logging frequency, position, context, and cited source per prompt, then roll up into a relative citation share of voice.
Most "AI visibility" content answers a different question than the one competitive teams actually have. Mapping which sources show up across ChatGPT, Perplexity, and AI Overviews for a topic tells you where the citation pool draws from — useful, but engine-centric. It doesn't tell you whether your competitor is beating you on the exact questions your buyers are asking, or why.
This is a narrower, more useful exercise: pick a fixed set of prompts a buyer would plausibly type, run each one against your brand and two or three named competitors, and record — prompt by prompt — who gets cited, where, in what context, and with which sources backing them up. It's a head-to-head study, not a source map. Run consistently, it turns "we think we're behind on AI search" into a list of specific prompts you're losing and specific reasons why.
Why the comparison has to be locked, not casual
Direct answer: A single ChatGPT query proves nothing. Model responses vary session to session, even with identical wording, because of retrieval variance and non-deterministic generation — which is why any credible benchmarking methodology insists on running each prompt multiple times across separate sessions before recording a result, and treating a one-off screenshot as evidence is the most common mistake teams make when they start this exercise (MaxAEO's benchmarking framework).
The two variables that have to stay fixed across the whole study are the prompt set and the engine set. If you test five prompts on ChatGPT for your brand and eight different prompts on Perplexity for a competitor, you don't have a benchmark — you have two unrelated anecdotes. Everything else (which competitor wins, which sources get cited, how position shifts) is only meaningful relative to that constant.
Building the shared prompt set
Direct answer: Don't start from your own keyword list. Start by running unbranded, category-level prompts and letting the engines tell you who they consider your competitive set — it's common for this to surface an adjacent vendor or a legacy competitor that no analyst report currently lists, simply because that competitor has old content the model still retrieves and trusts (Menra's competitive benchmarking primer). Assuming your competitor set instead of discovering it is the single most common way these studies produce misleading results.
Once the competitor set is confirmed, build 30–50 prompts (more if budget allows) that mirror how a buyer actually moves through a decision, not just head terms:
- Category/definitional — "what is [category]," "how does [category] work"
- Comparison — "[Brand] vs [Competitor]," "[Competitor] alternatives"
- Use-case — "best [category] for [industry/company size]"
- Objection/evaluation — "is [Competitor] worth it," "[Category] pricing"
Skip prompts that only mention your own brand by name — they flatter you and measure nothing competitive. The prompts that matter are the ones where the model has to choose who to surface, because that's the only place a real gap shows up.
What to log for every prompt × competitor pair
Direct answer: For each prompt, run it against every brand in the set on the same day, across the same engines, and record four things per mention:
| Signal | What it captures |
|---|---|
| Citation frequency | Did the brand appear at all, across N repeated runs of this prompt? |
| Position/prominence | First mentioned, buried in a list, footnoted as a source link only |
| Context/framing | Praised, neutral, caveated ("expensive but full-featured"), or absent |
| Cited source | Whose page, review, or third-party writeup backed the mention |
Roll these up per prompt into a simple win/loss call, and roll the whole set up into a share-of-voice number: your mentions divided by total mentions across the tracked competitor set for that prompt group. This is the same logic behind "citation share of voice" as a benchmarking metric — it's relative, not absolute, so a 20% share means something different in a five-competitor category than a two-competitor one.
The source column is the one teams skip and shouldn't. When a competitor wins a prompt, the citation backing them is rarely their own homepage — third-party coverage (reviews, comparison sites, press, structured Q&A pages) accounts for the large majority of what gets cited across models, with a brand's own site typically supplying only a small fraction of total citations (Ahrefs' 75,000-brand study). If a competitor is consistently winning a prompt off a third-party page, that's a PR/outreach problem, not a content problem on your site.
What tends to separate a winning citation from a losing one
Direct answer: Once you have a stack of prompt-level results, the useful next step is pattern-spotting: pull the transcripts for prompts your competitors won and look for what their cited content has in common. A few patterns show up repeatedly across independent research, though all of it should be read as correlation, not a guaranteed formula:
- Direct, front-loaded answers. Content that states the answer in the first paragraph and expands with evidence afterward tends to out-cite content that builds up to the point.
- Quotations and statistics. The Princeton/IIT-Delhi paper that coined the term Generative Engine Optimization tested nine content-level interventions on a 10,000-query benchmark and found that adding quotations from authoritative sources and adding relevant statistics produced the largest visibility lifts of the tactics tested, while keyword stuffing was the only tactic that measurably hurt (Aggarwal et al., "GEO: Generative Engine Optimization," arXiv).
- Off-site brand signal, not link volume. Across Ahrefs' analysis of 75,000 brands, branded web mentions correlated with AI-answer visibility roughly three times more strongly than backlink counts did — a signal about how much the brand is talked about, independent of who links to it (Ahrefs; follow-up analysis: Ahrefs).
- Technical baseline, not a technical hack. Semrush's analysis of five million cited URLs found that pages models cite tend to share strong technical foundations — but framed this explicitly as a correlation observed at scale, not a causal lever any single page can pull (Semrush).
- Consistency across studies matters more than any one number. A synthesis of 54 separate studies and experiments on AI citation factors scored each proposed factor by how repeatable it was across independent research — useful because it flags which "content patterns" are one lucky study versus a pattern that keeps showing up (Cyrus Shepard's factor analysis, via PPC Land).
When you compare your losing prompts against a competitor's winning content, check for these patterns specifically — a quoted stat from a named source, a direct answer in the first two sentences, a structured comparison table — rather than assuming length, keyword density, or schema markup alone explains the gap.
Say "correlated," not "caused" — and mean it
This is the part worth being disciplined about, because it's easy for a benchmarking exercise to slide into "competitor X does Y, therefore Y works." Every study cited above stops short of that claim, and for good reason. Seer Interactive's analysis of AI Overview citations and click-through rate found that cited brands saw meaningfully higher organic and paid CTR than uncited brands on the same queries — but the study's lead researcher was explicit about the limits: "We cannot definitively prove that citation causes higher CTRs; it's equally possible that brands with stronger authority and higher baseline CTRs are simply more likely to be cited by Google's AI" (Seer Interactive). Ahrefs' own writeup of its 75,000-brand correlation study carries the same caveat, and Semrush's technical-SEO study states it directly: these are correlations observed at scale, not causal proof for any individual page.
The same logic applies to your own prompt-by-prompt study, at smaller scale and with less statistical power. If a competitor wins six of ten prompts and their cited pages all happen to include a named statistic, that's a real, worth-testing pattern — not a confirmed rule. Treat what you find as a prioritized hypothesis list for content and PR experiments, not a checklist guaranteed to flip a citation in your favor. The honest version of this exercise produces a ranked set of things worth trying, backed by your own observed data, rather than a false promise that any one tactic will move a specific prompt.
Making it repeatable
The value of this method compounds with repetition — a single pass tells you where you stand today; the same prompt set run monthly tells you whether a content or PR push actually moved a specific prompt from a loss to a win. That's the operational reason to keep the prompt set, competitor set, and engine set locked: without that consistency, you can't tell a real shift from ordinary model-to-model variance. nqzai's own approach to this is to run a locked prompt set against your brand and named competitors on a recurring schedule, log citation position and source per prompt automatically, and flag week-over-week changes — so the comparison work described above happens continuously instead of as a one-time audit.
Sources:
- Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735
- Ahrefs — An Analysis of AI Overview Brand Visibility Factors (75K Brands)
- Ahrefs — Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews
- Semrush — How Do Technical SEO Factors Impact AI Search?
- Cyrus Shepard's AI citation ranking factors analysis, via PPC Land
- Seer Interactive — AI Overview Impact on CTR
- MaxAEO — AI Search Visibility Benchmarking: A Practical Measurement Framework
- Menra — What Is Competitive Benchmarking in AI Visibility?



