TL;DR
A September 2025 Search Engine Land experiment found that only a page with complete schema appeared in a Google AI Overview, yet a larger Ahrefs study of 1,885 pages found no clear positive effect on AI citations—and a small but statistically significant decline. The Princeton-led GEO paper (KDD 2024) tested 10,000 queries and identified adding statistics, citations, and quotations as among the strongest levers for improving whether a source is included in generative answers. A legitimate AEO engagement replaces keyword lists with an evidence plan: for each priority topic, an inventory of verifiable claims and proof, flagging where none exists. It also requires a technical audit that checks GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot individually, since many sites block these unknowingly. The real deliverable is a documented production pipeline with briefs, named reviewers, fact-checking against the evidence plan, and a QA pass—not a promise of citation counts.
The bottom line: a credible agency focuses on process and evidence you can audit
Most "AEO agency" pitches read like SEO proposals with a find-and-replace: swap "keywords" for "prompts," add a slide about ChatGPT, keep the same deliverables list. That's not a scope — it's a rebrand. A credible answer-engine-optimization engagement looks different in structure, not just vocabulary, because the target is different: instead of optimizing for a ranking algorithm you can query and observe (Search Console, rank trackers, SERP features), you're optimizing for a set of black-box generative systems that don't publish how they select, weight, or cite sources, and that change without notice.
That opacity doesn't mean AEO work is unmeasurable or unscopeable. It means the deliverables have to be built around evidence and process rather than promised outcomes. Below is what a credible AEO/GEO engagement should actually produce, organized the way a buyer should evaluate a proposal: technical foundation, content evidence, production operations, measurement, and — critically — a clear statement of what's out of anyone's control.
Layer 1: A technical audit you could hand to an engineer
Direct answer: Traditional SEO audits already established the bar here, and it carries over directly: a real audit is a working document with specific issues, severity, evidence, and an owner attached — not a slide deck of traffic charts. The same standard applies to AEO, with three areas that are AEO-specific enough to call out by name.
Crawlability for AI agents, not just search engines. Search and training crawlers are no longer the same handful of user-agents. GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, and PerplexityBot behave differently, and many sites are still running robots.txt and CDN/WAF rules written for the 2023 "block everything unfamiliar" panic — rules that silently exclude a site from citation eligibility on platforms it never meant to block. A real audit checks each of these bots individually, distinguishes retrieval/citation crawlers from training crawlers, and verifies the CDN layer matches what robots.txt actually says (the two drift more often than teams expect).
Structured data, scoped honestly. Schema is a legitimate deliverable, but the evidence on its payoff is more mixed than most agency pitches suggest, which is exactly why it belongs in the audit rather than the sales pitch. A September 2025 Search Engine Land controlled experiment found that of three near-identical pages, only the one with complete, well-implemented schema (Article + FAQPage + Breadcrumb) appeared in a Google AI Overview and achieved the best organic rank — while the page with no schema wasn't indexed at all. But a much larger, later Ahrefs study tracking 1,885 pages that added JSON-LD against 4,000 matched controls found no clear positive effect on AI citations, and a small but statistically significant decline in AI Overview appearances. Read together — as Search Engine Land's own follow-up piece on the hype gap notes — there are still no peer-reviewed studies isolating schema's causal effect on LLM citation behavior, and none of the major AI platforms disclose their indexing methods. The honest deliverable is a schema implementation plan scoped to indexation and correctness (fixing broken or incomplete markup, which is unambiguously bad) rather than a promise that schema will move citation counts.
Rendering and content access. If a page's substantive content only appears after client-side JavaScript execution, some crawlers and retrieval pipelines may not see it at all. The audit deliverable is a rendered-vs-source diff on template-level page types (not every URL), flagging where the answer-bearing content — the sentence or stat a model would actually quote — isn't present in the initial HTML.
Layer 2: An evidence plan, not a keyword list
Direct answer: The single clearest signal separating a real AEO program from a repackaged content calendar is whether the "content strategy" deliverable is a list of target prompts and queries, or an actual evidence plan: what claims the site can substantiate, with what proof, in what format.
This is grounded in the field's own founding research. The Princeton-led "GEO: Generative Engine Optimization" paper (Aggarwal et al., accepted at KDD 2024) — the paper that named this discipline — tested nine content-level optimization strategies across roughly 10,000 queries and found that adding statistics, citations, and quotations to a page were among the strongest levers for improving whether a source was included in a generative answer, with effects varying meaningfully by domain. That's a different exercise than keyword research: it means an agency should be producing, for each priority topic, an inventory of what verifiable evidence (data, named sources, direct quotes, first-party numbers) the client can put on the page — and flagging where none exists, which is itself a useful finding.
A credible evidence-plan deliverable includes:
- A prioritized topic/question map tied to buyer-relevant queries (not just high-volume search terms)
- For each topic, the specific claims the content will make and the evidence backing each one
- Format guidance suited to extraction — direct-answer openings, comparison tables, defined terms — because a generative system needs to lift a self-contained fragment, not infer meaning from a whole page
- Identification of content gaps where the client has no defensible evidence yet (a legitimate reason to delay publishing, not a reason to fabricate)
Layer 3: Content production and QA operations
Direct answer: This is where AEO work most resembles disciplined SEO or technical writing operations, and where the "deliverables" most often go missing in thin proposals. Production should be a documented pipeline, not a black box: briefs before drafts, named reviewers, fact-checking against the evidence plan built in Layer 2, and a QA pass confirming schema, internal links, and citations actually shipped as specified. If a vendor can't show you a sample brief-to-published artifact trail, "content production" is not actually a scoped deliverable — it's a placeholder.
Freshness and maintenance belong here too. Generative engines re-crawl and re-synthesize regularly, so a page that was accurate and well-cited at launch can go stale — pricing changes, discontinued claims, outdated statistics — in ways that quietly erode citation eligibility. A real content-ops deliverable includes a review cadence and an owner for updates, not just a one-time publish.
Layer 4: Measurement and reporting that doesn't invent precision
Direct answer: AI visibility measurement is young enough that even the leading commentary is explicit about its limits: unlike traditional rank tracking, there is no public rank data for generative answers, responses are probabilistic rather than fixed, and platforms sample and re-generate answers differently run to run. A credible reporting deliverable acknowledges this rather than presenting synthetic dashboards as if they were Search Console data.
That said, a defensible measurement practice does exist and should be part of the scope:
- Citation/mention tracking across a fixed, disclosed panel of prompts and platforms — sampled repeatedly over time, not a single snapshot, since any one run is noisy
- Share-of-voice by category, calculated as citation events won divided by total evaluations across the full prompt set (not an average of per-model percentages, which is a common and misleading shortcut)
- Multi-model reporting, since a single-platform reading (usually just ChatGPT) only describes that platform, and buyers query across several
- Clear separation between mentions (brand named, no link) and citations (attributed, linkable source) — HubSpot's practitioner guidance on AEO metrics treats this distinction as foundational, and reporting that blurs it overstates visibility
- A stated methodology: which prompts, which platforms, what sampling frequency, and what counts as a "win" — published alongside the numbers, not held back as proprietary magic
The reporting cadence matters as much as the metric definitions. A single month's snapshot is close to meaningless given how noisy generative outputs are; a trend across a quarter is what actually indicates whether a program is working.
| Deliverable area | What "credible" looks like | What "thin" looks like |
|---|---|---|
| Technical audit | Per-URL issues, severity, evidence, dev-ready tickets, AI-crawler-specific checks | A PDF of traffic charts and generic recommendations |
| Structured data | Scoped to indexation/correctness fixes, causal uncertainty disclosed | Sold as a guaranteed citation booster |
| Content strategy | Evidence plan: claims mapped to proof, gaps flagged | A keyword/prompt list relabeled "AEO strategy" |
| Content production | Documented brief→draft→fact-check→QA pipeline with named owners | "We'll write content for you," no visible process |
| Measurement | Multi-model, multi-prompt, trended, methodology disclosed | Single-platform snapshot presented as a score |
| Contract terms | Named limits on what can't be promised | Guaranteed rankings, citations, or "#1 in ChatGPT" |
Layer 5: What no agency — credible or not — can promise
Direct answer: This is the deliverable that's easiest to skip and most important to insist on: an explicit, written statement of what the engagement cannot guarantee, and why.
Google's own guidance is unambiguous on the search side, and the logic extends directly to AI answer engines, which are even less transparent. Google's "Do You Need an SEO?" documentation states plainly: "No one can guarantee a #1 ranking on Google. Beware of SEOs that claim to guarantee rankings, allege a 'special relationship' with Google, or advertise a 'priority submit' to Google." Google's companion guidance on third-party SEO tools adds that such tools and services "don't have access to Google's internal ranking data" and "can't guarantee performance." As Search Engine Journal reported, Google has since extended this same posture explicitly to cover AEO and GEO claims — reinforcing that the underlying systems, whether Google's ranking algorithm or an LLM's answer-generation process, are controlled by a third party outside any agency's or tool's reach. Search Engine Land's coverage of that update frames it the same way: check any agency's recommendations against primary guidance, and treat guaranteed-outcome claims as a red flag rather than a selling point.
For AI answer engines specifically, the uncertainty compounds. No major generative platform publishes its retrieval, ranking, or citation-selection logic. Outputs are probabilistic — the same prompt can return different citations on different runs. A credible agency puts this in writing as part of the deliverable, not the fine print: no one can guarantee inclusion in a specific AI answer, a specific citation count, or a specific share-of-voice number, because no one — including the platforms' own product teams in many documented cases — has full visibility into why a given source got cited on a given run.
What to actually ask for in a proposal
Direct answer: Before signing anything, ask for four artifacts, not four promises: a sample technical audit finding with severity and evidence attached; the evidence-plan template used for content topics (not a keyword list); a sample measurement report showing methodology and multi-model breakdown; and the written limitations clause. An agency that can produce all four is scoping a real discipline. An agency that substitutes a guarantee for any one of them is selling the thing Google's own guidance tells you to be wary of — just with "AI search" in place of "Google."
All of this assumes the decision to hire an agency has already been made. If that's still an open question, see our in-house vs. agency AEO decision framework for the six variables — including the technical-ownership bottleneck that determines whether either model can actually ship — worth working through first.
Sources:
- Do You Need an SEO? — Google Search Central
- Google's Guidance on Third-Party SEO Tools & Advice — Google Search Central
- Google's New Guidance Claims Authority Over SEO, Tools, And AEO/GEO — Search Engine Journal
- Google adds guidance on third-party SEO tools, services, advice — Search Engine Land
- GEO: Generative Engine Optimization — Aggarwal et al., arXiv:2311.09735 (KDD 2024)
- Schema and AI Overviews: Does structured data improve visibility? — Search Engine Land
- We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved. — Ahrefs
- How schema markup fits into AI search — without the hype — Search Engine Land
- AEO metrics every marketer should track — HubSpot
Evidence and scope
Review date: 2026-09-12.
Reproducible use. Use the framework with a defined audience, source data, and review date; test material recommendations against your own evidence before making a production or buying decision.
Limit. This article is educational guidance, not legal, financial, security, or performance assurance.



