TL;DR
In ChatGPT's B2B SaaS recommendations, G2 and Capterra each received zero citations across 233 recommendations, while independent blogs and vendor comparison pages captured 81.9% of third-party citations. Review platforms lost between 76.5% and 92.2% of organic traffic from early 2024 to late 2025, showing being cited and being visited have decoupled. Adding citations, direct quotes, and statistics to a page increased visibility in generative-engine answers by more than 40%,
A comparison page for AI search is a piece of content that answers "X vs Y" (or "X vs Y vs Z") with claims specific enough to be checked, sourced enough to be trusted, and structured enough for a generative engine to lift a paragraph and attribute it without rewriting it. That's the whole definition. Everything else — the table, the verdict, the tone — is in service of those three properties: checkable, sourced, extractable.
Most comparison pages fail at least one of them. That's not a stylistic problem anymore. It's a citation problem, and there's now a body of evidence showing AI answer engines are actively routing around the comparisons that fail.
The evidence: AI engines are discounting the pages that used to work
Three separate findings, from three independent studies, point the same direction.
First, review platforms are losing the citation share their traffic numbers would suggest they still have. SE Ranking's January 2026 analysis found that while G2, Capterra, Gartner Peer Insights, Software Advice, and TrustRadius together account for 88% of review-platform links inside Google AI Overviews, those same five platforms lost between 76.5% and 92.2% of their organic traffic between early 2024 and the end of 2025 — G2 down 84.5%, Capterra down 89%, TrustRadius down 92.2% (SE Ranking, "Review Platforms in AI Overviews," Jan 29, 2026). Being cited and being visited have decoupled.
Second, and more directly relevant to this piece: when researchers looked at what ChatGPT actually cites for B2B software recommendations — as opposed to what it names — the review platforms nearly vanish. DerivateX ran 40 B2B SaaS buying categories through ChatGPT ten times each (233 recommendations, 219 distinct tools) and pulled every cited source. Review aggregators accounted for 0.9% of citations total; G2 and Capterra each got zero. The vendor's own site was cited only 11.6% of the time. The other 88% of citations went to third parties — and 81.9% of those went to independent blogs and vendor-published comparison content, not major media, not Reddit. The study calls the gap between being recommended and being cited the "citation ownership gap" (DerivateX, "B2B SaaS AI Citation Study"). Every cited page in the sample used list structure; 68% included a comparison table; 56% included an FAQ section; 78% carried the current year in the title.
Third, the mechanism behind this isn't a black box. The foundational academic work on this — Aggarwal et al.'s "GEO: Generative Engine Optimization," presented at KDD 2024 — ran a black-box optimization study across roughly 10,000 queries and found that adding citations, direct quotations, and statistics to a page increased its visibility in generative-engine answers by more than 40%, while keyword-stuffing tactics that work in traditional SEO actually decreased visibility (Aggarwal et al., arXiv:2311.09735). Generative engines don't rank pages — they extract passages. A page that gives the model nothing extractable (vague claims, no numbers, no attribution) gives it nothing to cite.
Layer the citation-pattern research on top and the picture sharpens further. Profound's analysis of citations from August 2024 through June 2025 found that across ChatGPT, Google AI Overviews, and Perplexity, third-party and community sources dominate, with Reddit alone making up 46.7% of Perplexity's top-10 sources (Profound, "AI Platform Citation Patterns," June 5, 2025). A June 2026 analysis of 183,015 ChatGPT citation references across five review platforms found Gartner Peer Insights and G2 together made up 77.3% of review-platform citations specifically — but again, review platforms are a small slice of the total citation pool, not the dominant one (Parse.gl, "Which review site does ChatGPT trust most?," June 8, 2026). And Ahrefs' updated 2026 analysis of 863,000 keywords and 4 million AI Overview URLs found that only 38% of cited pages ranked in Google's top 10 for the same query, down from 76% eight months earlier — meaning traditional rank is a weaker and weaker proxy for what actually gets pulled into an AI answer (Search Engine Journal, "Google AI Overview Citations From Top-Ranking Pages Drop Sharply," March 2, 2026).
Put together: the platforms buyers trust for social proof (G2, Capterra) are not the pages AI engines cite when explaining a recommendation. Independent, structured, evidence-dense comparison content is.
Honest comparisons vs. vendor-biased ones
The distinguishing signal isn't who publishes the page — a vendor can write an honest comparison, and an "independent" blog can write a biased one. It's whether the content is checkable and balanced. A December 2025 breakdown from Unusual lays out the pattern generative engines seem to reward: compare against your strongest competitor, not a weak one; state tradeoffs explicitly ("faster, but uses more memory") instead of one-sided wins; replace adjectives ("enterprise-grade," "seamless") with numbers a reader could verify independently; and give explicit decision logic — "choose X if you need Y," not "we're the best choice for everyone" (Unusual, "How to write comparison content that AI models trust," Dec 7, 2025). The mechanism matches the GEO paper's findings almost exactly: specificity is what makes a claim extractable and attributable, and extractability is what gets a passage quoted.
Comparison-page patterns and their citation-worthiness
| Pattern | What it looks like | Citation-worthiness | Why |
|---|---|---|---|
| Vendor "we win everywhere" page | Every category shows the publisher's product ahead, no losses conceded | Low | No tradeoffs to extract; reads as marketing, contradicted by other sources the model has seen |
| Bare-bones feature checklist | Two columns of checkmarks, no context, no numbers | Low | Nothing to attribute — a checkmark isn't a checkable claim |
| Review-platform aggregate page | Star ratings and review counts, no narrative reasoning | Medium | Trusted for social proof, but thin on the specific, quotable claims models pull for "why" answers |
| Independent blog roundup, undated | Ranked list, no methodology, no publish/update date | Medium-low | Structure helps, but staleness and lack of sourcing undercut trust once checked |
| Evidence-backed comparison with decision logic | Named competitors treated fairly, explicit tradeoffs, dated, sourced claims, "choose X if / choose Y if" guidance | High | Matches every factor the GEO paper and citation studies associate with extraction: specificity, structure, balance, freshness |
| First-party data comparison | Same as above, plus original testing, benchmarks, or usage data no one else has published | Highest | Adds information density beyond what the model can already synthesize elsewhere — the citation compounds because it's the only source for that fact |
How to build one: a 9-step process
- Pick a comparison a real buyer is actually deciding between, not a strawman. If nobody genuinely weighs your product against the one you're comparing it to, the page reads as manufactured — and both the studies above and basic reader trust penalize that.
- Name your strongest competitor, not your weakest. Comparing against an easy target is the single fastest way to get discounted; models cross-reference other sources and notice when a "competitor" is chosen for how bad it looks.
- Replace every adjective with a number or a fact. "Fast" becomes a benchmark. "Secure" becomes the specific certification or control. If a claim can't be attached to something checkable, cut it or footnote it.
- Concede at least one real weakness per product, including your own. This is the single highest-leverage move in the research above — acknowledged tradeoffs are what separate a comparison from an advertisement, for both AI extraction and human trust.
- Build a comparison table with a stable, consistent criteria set — same rows for every product, same units, same time period. 68% of the pages ChatGPT actually cited in the DerivateX study used one.
- Write explicit "choose X if / choose Y if" guidance instead of a single overall winner. This format matched what both the Unusual analysis and the DerivateX structural findings identified as citation-correlated.
- Cite your own sources inline — pricing pages, documentation, published benchmarks, dated screenshots. A comparison that asks the reader (or the model) to take its word for it is exactly the pattern AI engines are learning to discount.
- Date the page and actually update it. Comparison content ages within months as pricing and features change; a visible "last reviewed" date is a low-cost trust signal, and letting it go stale is one of the fastest ways to lose it.
- Publish a public, crawlable version — not a gated PDF. If the comparison a reader most needs is behind a form, neither buyers nor generative engines can retrieve it. The most-cited pages in the studies above were all directly crawlable.
What this doesn't guarantee
Direct answer: Following this structure earns a page the chance to be extracted — it does not buy a citation outright, and it's worth being direct about the limits.
It won't override a weak evidence base. A perfectly formatted comparison built on invented numbers still fails the first time a model cross-references it against another source; the DerivateX and GEO findings both point to specificity being valuable because it's checkable, which cuts both ways.
It won't guarantee you win the comparison. Honest comparisons sometimes conclude the other product is the right fit for a given buyer — that's the cost of the format working at all. A comparison written to only ever conclude "we win" isn't following this pattern regardless of how the table is formatted.
It won't stay current on its own. The freshness signal identified across these studies decays; a page built once and left untouched drifts back toward the low-citation end of the table above within a couple of quarters as pricing, features, and competitors change.
It won't insulate you from the platform's own volatility. Citation behavior is still moving fast — the drop in AI Overview citations tracking top-10 organic rank, from 76% to 38% in eight months per Ahrefs, is one example of how quickly the underlying selection logic shifts (Search Engine Journal, March 2, 2026). And it arrives against a backdrop of declining user trust in AI search generally — a mid-2026 survey of 1,008 U.S. consumers found the share who rate AI search more helpful than traditional search fell from 82% to 54% year over year, even as usage kept rising (Digital Applied, citing the Fractl/Search Engine Land study, June 20, 2026). A well-built comparison page is a bet on a moving target, not a fixed rulebook.
Where nqzai fits
Direct answer: Building a comparison page that survives cross-referencing means knowing what's actually true about the competitor you're naming — their current pricing, their documented feature set, what independent sources say about them — not just what's true about your own product. nqzai's research and content tooling pulls and verifies that competitive detail directly from public sources at the time a comparison page is drafted, and flags claims that can't be backed by something checkable before they go live, so the page starts from evidence rather than assumption.
FAQ
Does a comparison page need to declare a winner to get cited?
No — the research points the other way. Pages with explicit "choose X if / choose Y if" segmentation performed better in the structural analysis of what ChatGPT actually cites than pages declaring a single universal winner, because segmented guidance maps to how buyers actually decide.
Is it worth still investing in G2 and Capterra profiles?
Yes, but for a different job. The DerivateX study found review platforms contribute almost nothing to ChatGPT's citation choices for B2B software (0.9% combined) even though the same buyers rely on them for social proof before ever asking an AI assistant. Treat review-platform presence as consideration-set validation, and independent comparison content as the citation layer.
How often does a comparison page need updating to stay citable?
There's no universal cadence in the research, but every study above ties trust and citation likelihood to visible freshness — a dated "last reviewed" stamp and periodic re-verification of pricing and feature claims. Quarterly is a reasonable floor for fast-moving categories.
Can a vendor write its own honest comparison, or does it have to come from a third party?
It can be vendor-published. The signal engines and researchers describe is specificity and balance in the content itself, not the domain it's hosted on — though a vendor page has to work harder to earn the same trust, since the default assumption is bias until the page demonstrates otherwise through conceded weaknesses and checkable claims.
What's the single fastest way to get a comparison page discounted?
Comparing against a strawman competitor, or claiming a universal win with no conceded weaknesses. Both patterns are explicitly called out across the sources above as the signature of marketing copy rather than analysis, and both are easy for a model to catch by cross-referencing other sources.
Does adding a comparison table by itself help?
It helps, but only as one factor among several. In the DerivateX sample, 68% of cited pages had a comparison table — a strong correlate, not a guarantee — and it needs to sit alongside dated sources, conceded tradeoffs, and specific claims to do the work the research associates with citation.



