TL;DR
Quotations beat raw statistics as the most citable claim type in the Princeton/Georgia Tech GEO benchmark — a 43.5% relative visibility lift vs. 34.2% for added stats, though stats win in Law & Government, Debate, and Opinion queries while quotes win in People & Society and History.
A separate analysis of 3 million ChatGPT responses found definitional phrasing like "X is" or "X refers to" was cited nearly twice as often, and 78.4% of citation-linked questions were pulled from a heading followed immediately by its answer. Comparative claims fail hardest: 18 of
Most GEO advice treats "citability" as a single dial you can crank up — add a stat here, a citation there, watch your visibility climb. That framing misses something the underlying research actually shows: different types of claims behave differently when a generative engine decides what to lift into its answer. A raw number, a direct quote, a definition, and a head-to-head comparison aren't interchangeable — they get extracted through different mechanisms, they fail in different ways, and they need to be written differently to survive.
This is a guide to those differences, grounded in what's actually been published, not in the kind of unverifiable multiplier stats that circulate in GEO content (you've probably seen the "X is 2x more likely to be cited" genre — most of it can't be traced to a real source, so it's left out here entirely).
Quick Answer
- If you're writing about People & Society or History → prioritize quotation-based claims, because the GEO-BENCH paper found quotations performed best in those domains, with a 43.5% relative visibility lift over baseline.
- If you're writing about Law & Government, Debate, or Opinion → prioritize raw statistics, because the same benchmark showed statistics performed best in those query types, with a 34.2% relative gain.
- If you're writing a definitional explainer → structure it as "X is" or "X refers to" directly under a heading stating the question, because Kevin Indig's study of 3 million ChatGPT responses found such phrasing was cited nearly twice as often as non-cited content.
- If you're writing a comparative claim → base it on named entities with a stated basis and measured difference (or use a table), because 18 of 21 participants in the Salesforce AI Research study reported missing citations for claims, and comparative/superlative claims without atomic facts are disproportionately exposed to that failure mode.
The four claim types
Direct answer: For content-strategy purposes, almost every citable sentence in a piece of B2B content falls into one of four buckets:
- Raw statistic — a number, usually with a unit and a timeframe ("SaaS churn averages 5–7% annually")
- Quote / attribution — words assigned to a named person or organization ("According to Gartner analyst X...")
- Definitional statement — a plain-language explanation of what a term or concept means
- Comparative claim — a statement that ranks or contrasts two or more things
Each has a different job in an argument, and — this is the part most page-structure advice skips — a different extraction profile inside an answer engine.
What the research actually shows
Direct answer: The most rigorous public source on this is still the Princeton/Georgia Tech "GEO: Generative Engine Optimization" paper, first posted in late 2023 and later published at KDD 2024 (arXiv:2311.09735; ACM DOI). The authors built a 10,000-query benchmark (GEO-BENCH) across eight domains and tested nine content-modification strategies against a simulated generative engine calibrated to Bing Chat, then spot-checked the strongest tactics on Perplexity.
Two of those nine strategies map directly onto claim types here, and the paper reports specific relative gains over baseline on its "Position-Adjusted Word Count" visibility metric:
- Quotation Addition produced the largest lift of any tested method — roughly a 43.5% relative improvement over baseline
- Statistics Addition produced roughly a 34.2% relative improvement
- Cite Sources (adding attributions/citations to claims) produced roughly a 29.0% relative improvement
That ordering is worth sitting with, because it cuts against the common assumption that numbers are the top-performing claim type. In this benchmark, well-sourced quotations outperformed added statistics. The paper also found the effect is domain-dependent, not universal: statistics performed best in "Law & Government," "Debate," and "Opinion" query types, while quotations performed best in "People & Society," "Explanation," and "History." Authoritative-tone rewrites did best in "Debate," "History," and "Science." In other words, the right claim type to lean on depends on what kind of question you're answering — a benchmark comparison post and an explainer post should not be built around the same claim mix.
One important caveat: this is one benchmark, on one simulated engine, measuring one visibility metric. Treat the ordering as directional evidence about which claim types are more extractable, not as a fixed multiplier you can apply to your own content.
Definitions have their own signature
Statistics and quotes were directly tested in the GEO paper; definitions weren't one of its nine strategies. But a separate, large-scale analysis gives a clear, independently-sourced picture of how definitional language behaves. Growth researcher Kevin Indig's study of roughly 3 million ChatGPT responses and 18,000+ verified citations, published via Search Engine Land, found that cited passages used explicit definitional phrasing — constructions like "X is" or "X refers to" — nearly twice as often as non-cited content from the same pages. The same study found cited passages were roughly twice as likely to contain a question mark, and that 78.4% of citations tied to a question were pulled from a heading followed immediately by its answer — consistent with models treating an H2 as a prompt and the paragraph beneath it as the answer.
The practical read: a definitional claim gets extracted when it's phrased as a plain, self-contained sentence sitting directly under a heading that states the question it answers — not when the definition is buried three sentences into a paragraph that opens with throat-clearing.
Comparative claims are the hardest to make citable
Comparisons are where things get genuinely difficult, and the research explains why. A 2025 study from Salesforce AI Research, published at the ACM Conference on Fairness, Accountability, and Transparency, had 21 participants use answer engines (Perplexity, You.com, Bing Chat) against traditional search and catalog failure modes. "Missing citation for claims and information generated" was reported by 18 of 21 participants — one of the most common complaints in the study (ACM DOI). An earlier foundational benchmark, "Evaluating Verifiability in Generative Search Engines" (Liu, Zhang & Liang, Findings of EMNLP 2023), found that across several commercial answer engines, a large share of generated statements weren't fully supported by the citations attached to them (ACL Anthology; arXiv:2304.09848).
Comparative and superlative claims are disproportionately exposed to this failure mode, because "X is better than Y" or "X is the fastest" is exactly the kind of statement a model can generate without needing to lift it from a specific source — there's no atomic fact to extract and attribute, just a judgment to reproduce. The fix isn't to avoid comparisons; it's to make them checkable. A comparison built from named entities, a stated basis of comparison, and a specific measured difference gives the model something concrete to quote and attribute. A comparison that's just an adjective ("more scalable," "industry-leading") gives it nothing to lift, so it either gets paraphrased with no citation or dropped.
Several independent practitioner analyses also report that comparisons presented as structured tables get pulled into AI answers more consistently than the same comparison written out in prose — the logic being that a table row already states the relationship between two facts explicitly, while a paragraph requires the model to infer that relationship from surrounding sentences. That pattern is directionally consistent with what's published, but the specific multiplier figures attached to it in various blog posts aren't traceable to a primary study, so treat "tables help" as a reasonable structural bet, not a proven ratio.
Claim-type cheat sheet
| Claim type | What makes it extractable | Where it performs best (per GEO-BENCH domains) | How to write it |
|---|---|---|---|
| Raw statistic | Self-contained number + unit + timeframe + source | Law & Government, Debate, Opinion | State the number in one sentence, name the source and date inline — don't make the reader infer the source from a footnote |
| Quote / attribution | Named speaker, short, unambiguous | People & Society, Explanation, History | Attribute by name and title in the same sentence as the quote; keep it under ~30 words |
| Definitional statement | Plain "X is / X refers to" phrasing, sitting under a heading | Explanation-heavy, "what is" queries | Ask the question as an H2, answer it in the very next sentence, no preamble |
| Comparative claim | Named entities + explicit basis + measurable difference | Debate, commercial "X vs. Y" queries | Replace adjectives ("better," "leading") with a stated axis and a number or named criterion; consider a table |
Auditing your own content
Direct answer: If you're rewriting existing pages for citability, the highest-value pass isn't a keyword check — it's a claim-type check. For every factual sentence on the page, ask:
- Is this a statistic without a source or date attached? Add both inline; an unsourced number is one of the easiest things for a model to paraphrase and strip attribution from.
- Is this a quote buried mid-paragraph instead of clearly attributed? Pull it out, name the speaker, keep it short.
- Is your definition three sentences deep instead of sitting right under its own heading? Move it to the first sentence after the H2.
- Is your comparison an adjective instead of a measurement? Replace "faster" or "more comprehensive" with the actual axis of comparison and, where you have one, a real number.
None of this requires restructuring your whole page — the pattern that keeps showing up across this research is that extractability is mostly a sentence-level property, not a page-level one. A page can have excellent structure and still bury its best claims in unextractable phrasing, and a fairly plain page can out-cite it by making every factual sentence stand on its own.
Sources:
- GEO: Generative Engine Optimization (Aggarwal et al., arXiv:2311.09735)
- GEO: Generative Engine Optimization, KDD 2024 (ACM DOI)
- ChatGPT citation location & language study (Kevin Indig, Search Engine Land)
- Search Engines in the AI Era: A Qualitative Understanding to the False Promise of Factual and Verifiable Source-Cited Responses in LLM-based Search (ACM FAccT 2025)
- Evaluating Verifiability in Generative Search Engines (Liu, Zhang & Liang, Findings of EMNLP 2023)
- Evaluating Verifiability in Generative Search Engines (arXiv:2304.09848)