---
title: "How to Vet an AEO Agency: 5 Red Flags Before You Sign"
description: "A buyer's guide to spotting AEO/GEO agency red flags — guaranteed citations, opaque methods, fabricated benchmarks, weak reporting, and scope that hides dependencies."
answer_summary: "A buyer's guide to spotting AEO/GEO agency red flags — guaranteed citations, opaque methods, fabricated benchmarks, weak reporting, and scope that hides dependencies."
canonical: "https://nqz.ai/blog/geo-aeo-agency-red-flags-claims-reporting-and-scope-problems"
published_at: "2026-07-18T10:10:56.229Z"
updated_at: "2026-09-10T12:25:07.109Z"
author: "nqzai Editorial Team"
category: "GEO"
tags: ["AEO","GEO","agency vetting","AI search visibility","marketing procurement","answer engine optimization"]
image: "https://nqz.ai/blog/covers/geo-aeo-agency-red-flags-claims-reporting-and-scope-problems.webp"
---

# How to Vet an AEO Agency: 5 Red Flags Before You Sign

The market for AI-search visibility services exploded almost overnight. A discipline that didn't have a settled name eighteen months ago — call it AEO, GEO, AIO, or LLMO depending on which agency's deck you're reading — is now sold by hundreds of vendors, many of them former SEO shops that added "AI visibility" to their service list without changing much else underneath. That's not automatically a problem. But it does mean the market is unusually thin on standards, unusually thick on unverifiable claims, and unusually easy to oversell to a buyer who doesn't already know what "good" looks like.

The good news is that most of the red flags aren't new. They're the same patterns that have burned buyers of traditional SEO services for two decades, adapted to a new set of nouns. If you know what to listen for, a 30-minute vendor call will tell you most of what you need to know.

## Red flag 1: Guarantees of specific citations or rankings

**Direct answer:** Any agency that promises a specific outcome — "we'll get you cited in ChatGPT for your category," "guaranteed top-3 in AI Overviews," "guaranteed mentions within 90 days" — is promising something no vendor can control. This isn't a new argument; it's the same one Google has made about traditional rankings for years. Google's own developer documentation states plainly: "No one can guarantee a #1 ranking on Google," and it explicitly tells site owners to be wary of any SEO that claims to guarantee rankings ["Do You Need an SEO? Tips for Hiring an SEO," Google Search Central](https://developers.google.com/search/docs/fundamentals/do-i-need-seo).

Google recently extended that same document to cover AI-answer optimization directly, adding language asking buyers to check whether an SEO's "advice on optimizing for AI experiences (also known as 'AEO' or 'GEO' services)" is aligned with Google's own guidance, and warning that tactics used to chase visibility in generative answers can cross the line into spam. For the first time, the page also directs businesses toward the FTC's complaint process if they believe they've been sold fraudulent SEO or AI-optimization services [Search Engine Land, "Google adds guidance on third-party SEO tools, services, advice"](https://searchengineland.com/google-adds-guidance-on-third-party-seo-tools-services-advice-and-updates-hiring-an-seo-doc-479637); [Search Engine Journal, "Google's Updated Guidance Urges FTC Complaints Against Shady SEOs"](https://www.searchenginejournal.com/googles-updated-guidance-urges-ftc-complaints-against-shady-seos/578151/).

The logic extends cleanly to AI engines. No agency has a commercial relationship with an LLM provider that lets it dictate which sources get cited in a generated answer. Model weights, retrieval behavior, and synthesis logic change on a schedule the agency doesn't control and often can't observe. An agency can influence the inputs — content structure, entity clarity, source credibility signals — but it cannot guarantee the output, any more than an SEO could guarantee a #1 organic ranking. Treat "guaranteed citations" exactly like "guaranteed #1 ranking": a claim that only makes sense if the vendor is either targeting outcomes so narrow they're meaningless, or planning to use tactics that create downside risk later.

## Red flag 2: Opaque or unexplained methodology

**Direct answer:** A second, related pattern: agencies that describe their process in outcome language ("we optimize your brand for AI visibility") but can't or won't explain the mechanism. Ask what specifically changes on the site or in the content, and you should get a concrete answer — structured markup, clearer entity definitions, content restructured to answer discrete questions, technical crawlability fixes, citation-worthy sourcing. If the answer stays abstract after two follow-up questions, that's the tell.

This maps directly onto Google's revised guidance, which asks buyers to check whether recommended tactics are "aligned with Google Search's official guidance on optimizing for generative AI features," and separately warns that third-party SEO/AEO tools aren't evaluated or endorsed by Google and don't have access to Google's internal ranking data — meaning their "AI visibility scores" are proprietary estimates, not ground truth [Google Search Central](https://developers.google.com/search/docs/fundamentals/do-i-need-seo). A legitimate agency will tell you what its tools measure and where the estimate comes from. A vendor selling a black box — "our proprietary AI visibility index," no further explanation offered — is asking you to trust a number it won't show its work on.

## Red flag 3: Fabricated or unverifiable case studies and benchmarks

**Direct answer:** edu/en/publications/geo-generative-engine-optimization/). That "40%" figure is a maximum observed under a synthetic, black-box evaluation setup — not an average result, and not a number that transfers automatically to a live client engagement.

The single most-cited piece of research in this industry is the Princeton/Georgia Tech/Allen Institute paper on Generative Engine Optimization, which found that tactics like adding statistics, quotations, and source citations could boost visibility in generated answers by "up to 40%" in a controlled benchmark [Princeton University, "GEO: Generative Engine Optimization"](https://collaborate.princeton.edu/en/publications/geo-generative-engine-optimization/). That "40%" figure is a maximum observed under a synthetic, black-box evaluation setup — not an average result, and not a number that transfers automatically to a live client engagement. It's routinely stripped of that context and repeated as if it were a guaranteed or typical client outcome. If an agency cites it (or any similar study) without the caveats, or presents its own "case studies" without a defined baseline, a stated measurement window, or a way for you to verify the underlying data, treat the claim as unsubstantiated until proven otherwise.

This isn't just a research-integrity nitpick — it's a compliance issue. Under the FTC's advertising substantiation doctrine, a company must have a reasonable basis for a claim *before* it's made, not after ["FTC Policy Statement Regarding Advertising Substantiation," Federal Trade Commission](https://www.ftc.gov/legal-library/browse/ftc-policy-statement-regarding-advertising-substantiation). The same standard applies to testimonials and case studies used in marketing: results shown must reflect what customers can generally expect, and if they don't, that has to be clearly disclosed [FTC Endorsement Guides FAQ](https://www.ftc.gov/business-guidance/resources/ftcs-endorsement-guides-what-people-are-asking); [16 CFR Part 255, eCFR](https://www.ecfr.gov/current/title-16/chapter-I/subchapter-B/part-255). Ask directly: is this case study representative, or a best-case outlier? Can I talk to the client? What was the baseline before the engagement started? A vendor that can't answer is asking you to buy on faith in a market where "faith" has already produced a lot of recycled, unverifiable statistics.

## Red flag 4: Weak or vague measurement and reporting

**Direct answer:** AI-answer visibility is genuinely harder to measure than classic SERP rank tracking — there's no universal "position 1" equivalent, engines don't expose consistent APIs for citation tracking, and results vary by query phrasing, user history, and model version. That difficulty is real, and it's exactly why vague reporting is such an easy place to hide. Watch for reports built entirely around soft, undefined terms — "AI visibility improved," "mentions trending up," "sentiment positive" — with no baseline, no defined set of tracked queries or engines, and no connection back to traffic, leads, or pipeline.

Good reporting names the engines and query set being tracked, shows a defined baseline and change over time, and — critically — ties the AI-visibility metric to a downstream business outcome the way legacy SEO reporting learned to tie rankings to traffic and traffic to conversions. If a reporting deck can't answer "compared to what, and so what," it's not measurement, it's a status update dressed as one.

## Red flag 5: Scope language that hides dependencies

The final pattern is contractual rather than analytical. Some AEO/GEO engagements are scoped so that nearly all the technical work — implementing structured data, fixing crawlability, restructuring content, maintaining any AI-crawler-facing configuration — sits with the client's own engineering or content team, while the agency's deliverable is strategy documents, audits, and monthly reporting. That can be a legitimate division of labor. It becomes a red flag when the scope language obscures it: when "full-service AI optimization" turns out to mean the agency diagnoses and the client executes, and any visibility gain gets attributed to the agency regardless of who did the work.

Before signing, get an explicit breakdown of who owns each deliverable — audit, strategy, content production, technical implementation, ongoing monitoring — and what happens to the timeline and the price if your team can't staff its side. A vendor that resists writing this down in the contract is often the same vendor whose case studies (see red flag 3) quietly credit itself for work the client actually did.

## Red flag vs. what good looks like

| Red flag | What good looks like |
|---|---|
| Guarantees specific citations, rankings, or a fixed number of AI mentions by a date | States realistic ranges, explains what it can and can't control, and points to process metrics it's actually accountable for |
| Describes tactics only in outcome language; won't explain the mechanism | Names concrete technical and content changes, and how they map to Google's and other engines' public guidance |
| Cites benchmarks or case studies without context, baseline, or verification | Provides baseline, methodology, and a client reference you can actually contact |
| Reports "visibility" or "sentiment" with no defined engines, queries, or baseline | Names the tracked engines/queries, shows a baseline and trend, ties the metric to traffic or pipeline |
| Scope hides who does the technical work; agency takes credit regardless | Contract specifies ownership by deliverable, with a RACI or equivalent breakdown |

## Practical due diligence

**Direct answer:** com/search/docs/fundamentals/do-i-need-seo)? Ask for a client reference you can call, not just a logo on a slide.

Before signing anything, ask the same questions Google's own hiring guidance recommends for traditional SEO, adapted to AI search: what work will happen, why does it matter, how will success be measured, and will the agency explain each change as it's made [Google Search Central](https://developers.google.com/search/docs/fundamentals/do-i-need-seo)? Ask for a client reference you can call, not just a logo on a slide. Ask what happens contractually if the promised range isn't hit. And if you've already been sold guarantees that didn't materialize, or case studies that don't check out under scrutiny, the FTC's complaint process exists for exactly this kind of deceptive marketing claim, and Google's guidance now points buyers toward it directly [Search Engine Journal](https://www.searchenginejournal.com/googles-updated-guidance-urges-ftc-complaints-against-shady-seos/578151/).

None of this means AEO/GEO work is snake oil — the underlying discipline of making content legible, well-sourced, and structured for machine synthesis is real and worth investing in. It means the market for selling that work hasn't caught up to the market for doing it, and until it does, the burden of verification sits with the buyer.

Sources:
- [Google Search Central — "Do You Need an SEO? Tips for Hiring an SEO"](https://developers.google.com/search/docs/fundamentals/do-i-need-seo)
- [Search Engine Land — "Google adds guidance on third-party SEO tools, services, advice and updates hiring an SEO doc"](https://searchengineland.com/google-adds-guidance-on-third-party-seo-tools-services-advice-and-updates-hiring-an-seo-doc-479637)
- [Search Engine Journal — "Google's Updated Guidance Urges FTC Complaints Against Shady SEOs"](https://www.searchenginejournal.com/googles-updated-guidance-urges-ftc-complaints-against-shady-seos/578151/)
- [FTC — "FTC's Endorsement Guides: What People Are Asking"](https://www.ftc.gov/business-guidance/resources/ftcs-endorsement-guides-what-people-are-asking)
- [FTC — Policy Statement Regarding Advertising Substantiation](https://www.ftc.gov/legal-library/browse/ftc-policy-statement-regarding-advertising-substantiation)
- [eCFR — 16 CFR Part 255, Guides Concerning Use of Endorsements and Testimonials in Advertising](https://www.ecfr.gov/current/title-16/chapter-I/subchapter-B/part-255)
- [Princeton University — "GEO: Generative Engine Optimization"](https://collaborate.princeton.edu/en/publications/geo-generative-engine-optimization/)
- [Shortlist — "Can an SEO Agency Guarantee Rankings? The Honest Answer"](https://shortlist.io/blog/can-an-seo-agency-guarantee-rankings/)
