TL;DR

91% of executives have asked about AI search visibility, yet 62% of practitioners say it drives less than 5% of measurable revenue — a gap a proper audit must close. Ahrefs tracked AI Overview citations from top-10 organic pages dropping from 76.1% in mid-2025 to just 38% by March 2026, proving that ranking high no longer guarantees AI citation. The same firm found AI Overviews now cost the top organic result 58% of its expected clicks, up from 34.5% a year earlier — making answer visibility the only visibility that still pays.

A credible GEO audit deliverable must measure baseline citation frequency per engine, competitive share of voice, technical crawlability, content extractability (answer-first framing, statistics, citations), entity signals, and AI-referral traffic — then connect every finding to a ranked action list and a way to prove movement on the next pass, or agencies will sell thin templates that get exposed the minute a client asks for ROI.

A GEO audit deliverable is a written assessment of how often, where, and in what form a brand appears in AI-generated answers — ChatGPT, Google's AI Overviews and AI Mode, Perplexity, and comparable systems — paired with a prioritized set of changes intended to improve that appearance. It is not an SEO audit with an "AI" section bolted on, and it is not a screenshot of a chatbot answering one prompt. Done properly, it functions the way a technical SEO audit has functioned for two decades: baseline measurement, gap analysis against competitors, a ranked action list, and a way to prove movement on the next pass.

The problem agencies are running into is that there's no shared standard for what "properly" means. A September 2025 survey of SEO practitioners run by Aleyda Solis and covered by Danny Goodwin at Search Engine Land found that practitioners can't even agree on what to call the discipline — 36% said "AI search optimization," 27% stuck with plain SEO terminology, and only 18% used "generative engine optimization." The same survey found 91% of respondents said leadership had asked about AI search visibility in the past year, but 62% said AI search was driving less than 5% of measurable revenue, with attribution cited as the main blocker. That gap — high executive demand, low measurement maturity — is exactly the space a well-built audit deliverable is supposed to fill, and exactly where a thin, templated one gets exposed.

Where the term actually comes from

"Generative Engine Optimization" isn't marketing copy invented by an agency. It's the title of a 2023 research paper — Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande, "GEO: Generative Engine Optimization," arXiv:2311.09735 — from researchers at Princeton, the Allen Institute for AI, Georgia Tech, and IIT Delhi, later published at KDD 2024. The paper introduced GEO-bench, a benchmark of real user queries, and tested specific content interventions (adding citations, quotations, statistics, and clearer structure) against a generative engine's tendency to reference a source. The headline finding — that these interventions could lift visibility in generated responses by up to 40% — is the number the industry has run with. It's worth reading the paper directly before quoting that figure, because the gain varied substantially by domain and query type, and the study measured a research system's behavior at a point in time, not a guarantee about any live production model today. An audit deliverable that cites this paper should say that plainly, not treat it as a universal formula.

What actually gets cited, and how unstable it is

Direct answer: Two independent measurement efforts help ground the "what to measure" question in something other than vendor claims. Profound's analysis of roughly 680 million citations across ChatGPT, Google AI Overviews, and Perplexity between August 2024 and June 2025 found the platforms don't source alike at all: Wikipedia accounts for 7.8% of ChatGPT's citations but only 0.6% of Google AI Overviews' citations, while Reddit makes up 6.6% of Perplexity's citations versus 1.8% of ChatGPT's. An audit built around one engine's behavior will misread another engine entirely.

Ahrefs has tracked the second half of that instability — how loosely AI citation tracks organic rank — and the number has moved fast. In a July 2025 study of 1.9 million citations, Ahrefs found 76.1% of AI Overview citations came from pages already ranking in the traditional top 10. By a March 2026 follow-up covering 863,000 keyword SERPs and 4 million AI Overview URLs, that figure had fallen to roughly 38%. Ahrefs attributes part of the shift to improved citation-parsing on their end and part to Google's "query fan-out" process, which pulls sources from a cluster of related sub-queries rather than the direct SERP. Either way, a client who was told in mid-2025 that ranking well is "basically the same thing" as being cited is now working from a stale model. Separately, Ahrefs' February 2026 clickthrough analysis found that when an AI Overview appears, the top-ranking organic page now loses about 58% of the clicks it would otherwise have earned — up from a 34.5% loss measured in April 2025. That statistic belongs in every audit's executive summary, because it's the business case for doing the work at all: visibility in the answer is increasingly the only visibility that pays.

The anatomy of a credible deliverable

Direct answer: Most agency "AI audits" fail for the same reason thin SEO audits used to fail — they list observations without connecting them to a decision. A deliverable earns its fee when each section answers three questions: what was measured, why a client should care, and what evidence supports the finding.

Deliverable sectionWhat's measuredWhy it mattersTypical format
Baseline citation visibilityFrequency of brand mentions/citations across a fixed query set, per AI engineEstablishes the starting point before any work beginsScorecard by engine and query category
Competitive share of voiceSame query set run against named competitorsShows relative position, not just presenceStacked bar or ranked table
Technical crawlabilityrobots.txt rules for AI crawlers, JS-rendering dependency, sitemap coverage, response codesContent that can't be crawled can't be cited, regardless of qualityPass/fail checklist with URLs
Content structure & extractabilityAnswer-first framing, passage-level clarity, heading hierarchy, presence of stats/citationsAI systems retrieve and cite passages, not whole pagesPage-by-page annotated list
Entity & authority signalsStructured data (schema.org), Wikipedia/Wikidata presence, third-party mentions, consistent NAP/entity namingGenerative engines lean on corroborating third-party sources, not brand-owned claimsEntity map + gap list
AI-referral traffic & attributionSessions/conversions from AI-platform referrers in analytics, distinguished from organicTies visibility to a number finance will recognizeTrend chart + defined UTM/referrer methodology
Answer accuracy & sentimentWhat AI engines actually say about the brand, correct or notBeing cited inaccurately can be worse than not being citedVerbatim quote log with source URLs
Prioritized roadmapFixes ranked by estimated effort and expected impactConverts findings into a work plan the client can approveTable with owner, effort, timeline

Structuring the audit: a step-by-step process

  1. Define the query set before touching any tool. Pull branded, category, comparison ("X vs Y"), and long-tail informational queries the client's actual buyers would plausibly type or speak into an AI assistant. A generic list of head terms produces a generic, unusable audit.
  2. Run the baseline across each target engine separately. Don't average across ChatGPT, AI Overviews, AI Mode, and Perplexity — the Profound citation-pattern data above shows they don't behave alike, so a single blended score hides where the real gaps are.
  3. Benchmark named competitors on the identical query set. A client needs to know not just whether they appear, but who is appearing instead of them and why (domain, page type, publish date).
  4. Audit technical crawlability. Check robots.txt rules for AI crawlers, confirm content isn't gated behind client-side rendering the crawler can't execute, and confirm sitemap and canonical hygiene. A 2025 academic study of robots.txt gatekeeping tracking reputable sites over time found AI-crawler blocking rose from roughly 23% in September 2023 to nearly 60% by May 2025 — enough sites now self-exclude that this step can't be assumed away.
  5. Audit content structure and extractability, page by page for the highest-priority query set: does the page answer the query directly and early, is it chunked into self-contained passages, does it carry the citations and specifics the Princeton paper's interventions tested.
  6. Audit entity and third-party authority signals — schema markup, Wikipedia/Wikidata presence, consistency of brand naming, and independent (non-owned) sources that corroborate claims.
  7. Instrument AI-referral tracking in the client's analytics before the audit closes, not after, so the next re-audit has a real trend line instead of another one-time snapshot.
  8. Translate every finding into a ranked action — effort, expected impact, and an owner — rather than a flat list of observations.
  9. Package it for two audiences at once: an executive summary a non-technical stakeholder can act on in five minutes, and an appendix with the raw query-by-query evidence a skeptical technical lead can verify line by line.

Limitations — what this doesn't guarantee

Direct answer: Be direct with clients about the edges of this work, because overselling it is what makes the "thin templated audit" reputation stick.

  • There is no stable ranking concept to report against. Traditional SEO has position 1 through 10 on a fixed page. AI answers cite zero to several sources, the count and order change response to response, and — per Ahrefs' own data above — the link between organic rank and AI citation moved by roughly 38 percentage points in eight months. A "score" from this audit is a snapshot, not a guaranteed trajectory.
  • Underlying models change without notice. Citation behavior shifted materially around Google's rollout of a new default model powering AI Overviews in early 2026, according to Ahrefs' March 2026 analysis. An audit's recommendations can go stale faster than a traditional SEO recommendation would.
  • Some platforms are close to a black box. Search Engine Land's 2025 survey found attribution and measurement were the top-cited obstacles for practitioners, not lack of tactics. Agencies can measure citation frequency and referral traffic; they generally cannot see a platform's internal ranking or selection logic, and any deliverable that implies otherwise is overpromising.
  • The Princeton GEO paper's own results varied by domain — the interventions that lifted visibility in one query category did not perform uniformly in others. A single "GEO score" that ignores this variance risks giving false confidence.
  • Correlation between a fix and a citation gain is hard to isolate. Multiple things change between audits — the client's content, competitors' content, and the model itself. Present improvement as directional evidence, not a controlled experiment.

Where nqzai fits

This is squarely inside what nqzai's tooling already produces, which is why it's a legitimate service line for agencies rather than a hand-wavy add-on. nqzai runs a client's brand and competitor names against a defined query set across multiple AI engines and returns an AI visibility score — how often the brand is mentioned or cited, and how that compares to named competitors on the same queries — alongside the standard SEO signals (rankings, technical health, backlink profile) in the same report, so an agency isn't stitching together a GEO section and an SEO section from two different tools. It can flag technical crawlability issues (crawler access, structured data gaps, entity consistency) and produce the query-by-query evidence table an audit's appendix needs. What it does not do, and what no tool honestly can: predict how a specific model update will change citation behavior next quarter, or guarantee that a fix will produce a citation gain, given the variance the Princeton paper itself documented. Agencies should present nqzai's output as the measurement and evidence layer of the audit — the baseline, the competitive comparison, the tracking over time — and keep the judgment calls about what to prioritize and how to phrase recommendations as their own.

FAQ

Is a GEO audit different from an SEO audit, or a section inside one?

They should ship as one integrated deliverable that shares a content foundation, but the sections measure different things — SEO audits technical health and rank position; GEO audits citation presence and answer accuracy across AI engines. Selling them as entirely separate products usually just means the client pays twice for overlapping crawl and content work.

How many queries should a baseline query set include?

Enough to cover branded, category, and comparison intents without becoming unreadable — most credible audits land in the 15–40 query range, weighted toward what the client's actual buyers would ask, not head terms picked for volume.

How often should an agency re-run the audit?

Quarterly is a reasonable default given how fast citation behavior has moved (see the 76%-to-38% shift above); monthly for competitive or fast-moving categories, since a single model update can change results materially.

Can an agency guarantee citation improvement as a contract term?

No — and any agency that does is overselling. The honest framing is a ranked action plan and measured before/after visibility, not a guaranteed outcome, given how much of the citation logic sits outside anyone's control.

Do AI Overviews and other AI engines cite the same sources?

Not reliably. Independent research has found Google's own AI Overviews and AI Mode products agree on cited URLs only a small fraction of the time, and citation-source composition (how much comes from forums, news, or brand sites) differs sharply by platform — which is why a single-engine audit understates the real gap.