TL;DR
Over 60% of AI search tools misattributed sources in a Tow Center study, with Grok-3 wrong 94% of the time and even the best performer, Perplexity, still erring on 37% of queries. Separately, an analysis of nearly 8,000 AI citations found 82.5% pointed to deeply nested pages rather than homepages, meaning engines pull individual claims from specific pages.
Meanwhile, 63% of B2B buyers used AI in a recent purchase, but 94% fact-checked its output, and 53% talked to a real peer—so vague, unverifiable case studies fail both extraction and trust checks. The bottom-line verdict: a citable case study must include a named customer, a specific before/after number with baseline, a measurement method, a date range, and a direct attributed quote; anything less is marketing copy that AI cannot reliably cite and buyers will not believe.
A case study is evidence when an AI answer engine can extract a specific, attributable claim from it and reuse that claim without needing to trust the page it came from. Everything else — the client logo, the warm quote about "great partnership," the unnamed "leading enterprise" — is marketing copy that happens to be shaped like a case study. Generative engines don't cite marketing copy well, because they have no way to check it. They cite claims they can verify, attribute, and quote.
That distinction matters more than most content teams realize, and it's not a matter of opinion. It's been measured.
Why AI answer engines are picky about proof
That finding lines up with a separate, less flattering piece of research about how AI search actually behaves once it's citing things. In March 2025, the Tow Center for Digital Journalism at Columbia published AI Search Has a Citation Problem, a study by Klaudia Jaźwińska and Aisvarya Chandrasekar that ran 1,600 queries against eight AI search tools — ChatGPT Search, Perplexity, Perplexity Pro, Gemini, DeepSeek Search, Grok-2, Grok-3, and Copilot — by feeding each tool a verbatim excerpt and asking it to identify the source. More than 60% of responses across the tools were wrong in some way; ChatGPT Search was incorrect 67% of the time, Grok-3 was incorrect 94% of the time, and even the best performer, Perplexity, was still wrong on 37% of queries. The tools were also bad at declining to answer — they guessed, confidently, rather than saying they didn't know.
Put those two studies together and the incentive is obvious. AI search systems are already prone to misattributing and garbling content. A case study that gives them a clean, quotable, specific, single-sentence claim to work with is far more likely to survive that process intact than one that buries its result in adjectives.
What makes a case study "evidence" instead of "marketing"
Direct answer: 5% of citations point to specific, deeply nested pages, not homepages or generic hub pages. Engines are pulling individual claims off individual pages, not crediting a domain in the abstract.
A Search Engine Land analysis of nearly 8,000 AI citations across ChatGPT, Gemini, Perplexity, and AI Overviews, published by James Allen in May 2025, found that different engines skew toward different source types — ChatGPT leans on Wikipedia and established authorities, AI Overviews leans on blog-style articles and Reddit — but the common thread across all four is that 82.5% of citations point to specific, deeply nested pages, not homepages or generic hub pages. Engines are pulling individual claims off individual pages, not crediting a domain in the abstract.
That's the same mechanic industry write-ups on the topic describe when they talk about case studies specifically. An analysis from LSEO on structuring case studies for AI citation argues that the single highest-leverage habit is making each result statement standalone — a sentence an AI system can lift out of the page and quote correctly without needing the surrounding paragraph for context. "Revenue grew" is not standalone. "Monthly recurring revenue grew from $340,000 to $612,000 over five months following a redesign of the onboarding flow" is.
The trust side of the equation matters just as much as the extraction mechanics. Buyers are already skeptical of anything a vendor says about itself, and that skepticism doesn't go away when AI is doing the summarizing. The TrustRadius 2026 B2B Buying Disconnect Report, based on 1,862 buyers and 444 vendors, found that 63% of buyers used AI somewhere in a recent purchase journey, but 94% of those buyers said they fact-check what the AI tells them at least some of the time — and 53% went further and talked to a real peer who had used the product. A Gartner sales survey published in March 2026 found 67% of B2B buyers now prefer a rep-free buying experience — which means the case study, not a salesperson, is frequently the thing doing the persuading, whether a human or a model is reading it.
Taken together, this is why vague case studies fail on two fronts at once: a model can't extract a clean claim from them, and a skeptical buyer wouldn't believe the claim even if it could.
Weak evidence vs. strong evidence
| Pattern | Weak (marketing fluff) | Strong (verifiable evidence) |
|---|---|---|
| Customer identity | "A leading SaaS company" | Named company, with role/title of the person quoted |
| Outcome claim | "Significantly improved efficiency" | "Cut manual QA review time from 6 hours to 45 minutes per release" |
| Baseline | Not stated | Explicit before/after numbers, both stated |
| Timeframe | Not stated, or "recently" | Specific date range ("March-August 2026") |
| Methodology | Not described | How the result was measured and by whom (e.g., "measured via internal ticket-resolution logs") |
| Attribution | Unattributed or house-written | Direct quote, attributed to a named person, ideally independently verifiable (LinkedIn, company site) |
| Corroboration | Self-published only | Linked to a third-party mention — review site, press coverage, public data — where one exists |
| Limitations | Implied universal success | States what didn't change or what conditions the result depended on |
| Format | Long narrative paragraph, PDF, or slide deck | Structured page: problem, method, named metric, quote, date — extractable as standalone sentences |
A step-by-step process for building citable case studies
- Pick outcomes you can attribute to a number, not an adjective. If the honest answer is "it felt faster," you don't have a case study yet — you have a testimonial. Go find the number before you write anything.
- Get the baseline in writing before the intervention, not after. A before/after claim without a documented "before" is a claim your own team can't defend, let alone an AI system.
- Name the customer and the person, with permission. An anonymized case study can still be well-written, but it can't be verified by anyone reading it — human or machine — which caps how much weight it can carry.
- State the measurement method in one sentence. "Measured using Q2 vs. Q3 support-ticket volume in our own helpdesk" is a methodology. "Customers report saving time" is not.
- Write the result as a standalone sentence. Draft the single sentence you'd want an AI answer engine to quote verbatim, with the number, the timeframe, and the entity all inside that one sentence — then build the narrative around it, not the other way around.
- Attribute every claim near the claim itself, not in a footnote at the bottom. Systems that parse pages linearly associate authority with proximity — a citation link at the end of a 1,200-word page does less work than a number backed by a linked source in the same sentence.
- Link out to independent corroboration where it exists. A press mention, a public review, a regulatory filing, an earnings call — anything outside your own domain that supports the claim gives an outside system a second, independent path to the same fact.
- Disclose scope and limits. State what the result doesn't cover — company size, industry, starting conditions — so the claim reads as calibrated rather than absolute. Calibrated claims survive scrutiny better than sweeping ones.
- Mark the page up as a reviewable claim, not just prose. Schema.org's Claim and ClaimReview types exist precisely so a page can state "here is a specific, checkable assertion" in a machine-readable way — useful even outside formal fact-checking contexts, because it forces the writer to isolate the checkable claim from the surrounding narrative.
Where nqzai fits
Direct answer: nqzai's content tooling is built to catch the difference between a case study that reads well and one that would actually survive fact-checking — flagging unattributed claims, missing baselines, and results stated as adjectives instead of numbers before a page goes live, and checking that each core claim is written as a standalone, attributable sentence rather than something that only makes sense buried in a paragraph. It doesn't decide whether an AI engine cites the page; nothing can promise that. It's built to make sure the page has done everything within your control to earn it.
FAQ
Direct answer: There's no controlled study measuring named-vs-anonymous case studies specifically, but the mechanism is well established: named, attributed claims are independently checkable and self-contained claims extract cleanly, both of which the GEO paper and the Search Engine Land citation analysis link to higher visibility and citation accuracy.
How many case studies do we need before this is worth doing?
One well-built, verifiable case study with a real number and a named customer outperforms a dozen vague ones. Depth and verifiability matter more than volume here.
Should we still get anonymized case studies approved by legal — is there a middle ground?
Yes — you can often get sign-off to publish the metric and industry without the company name, or to use a title and industry ("VP of Operations at a mid-market logistics company") instead of a full identity. It's weaker evidence than a named source, but still stronger than an unattributed adjective.
Does adding a comparison table or FAQ to the case study page help it get cited?
The Search Engine Land citation analysis and the GEO paper both point toward structured, information-dense formatting helping extraction, but structure is a delivery mechanism — it doesn't substitute for the underlying claim being specific and attributed. A well-formatted vague claim is still a vague claim.
Is it worth adding ClaimReview or Claim schema markup to a normal case study page?
It's not required, and most case studies aren't formal fact-checks, so ClaimReview specifically may be a mismatch. But the underlying discipline it enforces — isolating one checkable assertion per section — is worth doing in plain text even without the markup.
How often should a case study be updated or re-verified?
Whenever the underlying number could plausibly have changed — a new pricing tier, a platform migration, a churned customer — because a stale or now-false number that's still published is worse for credibility than no case study at all, both for a human fact-checker and for a model that has no way to know the claim went stale.
How we keep this honest
Every response nqzai's agent generates is automatically graded by an independent AI judge for accuracy and whether it invents information it can't back up. As of September 2026: sampled responses averaged a 82% quality score over the trailing 7 days (n=39), and our nightly regression suite — which re-runs the agent against a fixed set of real scenarios — passed at a ~93% rate over the last 14 nights. This is internal automated QA, not an independently audited or third-party benchmark; we publish it as a transparency signal, not a claim of perfection.



