TL;DR

No official measurement API exists for any AI search engine — no query-level data, no impression counts, making every vendor and DIY method a sampled estimate, not ground truth. Ahrefs' Brand Radar pulls prompts from 100 billion keywords and Google's "People Also Ask" rather than user-written queries, while Semrush's Prompt Tracking runs daily against 289 million prompts across four engines.

Even the same brand will score differently across tools because each samples different prompt sets at different times against non-deterministic models. The article's verdict: treat any AI visibility score as a directional trend, not a precise rank, and run your own fixed prompt set on a consistent schedule to get usable comparative data.

Brand visibility in AI search is measured by running a fixed set of representative prompts against ChatGPT, Perplexity, Gemini, and Google AI Overviews on a repeating schedule, then logging whether your brand is mentioned, whether it's cited with a link, and how it compares to named competitors in the same answers. There is no equivalent of Google Search Console for any of these engines — no query-level API, no official impression count — so every vendor and every DIY method is running a sample and reporting an estimate, not a ground truth.

That distinction matters more than most measurement guides admit. Below is what the actual methodology looks like, what the credible vendors in this space (Ahrefs, Otterly.AI, Semrush, Profound) disclose about how they collect data, what independent research says about how noisy that data really is, and a practical way to run this yourself without overstating precision you don't have.

Quick Answer

  • If you're a small team with no budget and full control over your prompt set → use a manual/DIY spreadsheet, because it's zero cost but doesn't scale past ~20 prompts.
  • If you need broad, real-demand coverage and don't mind less control over exact wording → use Ahrefs Brand Radar, because it auto-generates prompts from Google's "People Also Ask" and a corpus of over 100 billion keywords.
  • If you want daily monitoring with prompt-level control across ChatGPT, Google AI Overviews, Perplexity, and Copilot → use Otterly.AI, because it runs daily and lets you manually add or discover the prompts you want tracked.
  • If you need integration with your existing SEO rank tracking and daily runs against Google AI Mode, AI Overviews, Gemini, and ChatGPT Search → use Semrush AI Toolkit, because it uses the same crawling infrastructure as its traditional rank tracker against a claimed database of 289 million prompts.

What "AI visibility" actually means

Four terms get used loosely and interchangeably. They aren't the same thing:

  • Mention rate — the percentage of tracked prompts where your brand name shows up anywhere in the answer, cited or not.
  • Citation rate — the percentage of mentions that come with an actual link or named source attribution back to your domain, as opposed to the model just recalling your brand from training data.
  • Share of voice / share of model — your mention or citation count relative to named competitors across the same prompt set, not an absolute number in isolation.
  • Sentiment — whether the mention is framed positively, neutrally, or negatively, which recent research found is far less stable run-to-run than whether a brand is mentioned at all.

Semrush's own metrics documentation defines its AI Visibility Score as a 0–100 benchmark of "how often your brand appears in AI-generated answers compared to competitors," built from these same underlying counts — not a single proprietary signal (Semrush AI Visibility Metrics). Profound frames its core number the same way: a percentage of mentions out of total tracked responses, explicitly described as "a quick snapshot," not an authoritative rank (Profound: How to Track Your Brand's Visibility in AI Search).

How the measurement actually works

Direct answer: Strip away the dashboards and every credible approach — vendor or DIY — follows the same four steps:

  1. Build a prompt set. A list of realistic queries covering direct brand questions ("what is [brand]"), category/problem prompts ("best tool for X"), and head-to-head comparisons ("[brand] vs [competitor]").
  2. Run each prompt against each engine on a schedule. Manually, via browser automation, or via a vendor's infrastructure.
  3. Parse the raw response. Detect brand and competitor mentions, extract any cited URLs or named sources, and score sentiment.
  4. Aggregate into rates and trends. Mention rate, citation rate, and share of voice over time — not a single run.

The differences between vendors are mostly in step 1 (where the prompts come from) and step 2 (how often they run). Ahrefs' Brand Radar methodology post is unusually specific about this: it pulls queries from Google's "People Also Ask" corpus and Ahrefs' own keyword database (they cite over 100 billion keywords tracked) rather than writing synthetic prompts, on the reasoning that real search demand produces more representative phrasing than a marketer guessing what buyers ask. It refreshes its question set monthly and reports over a rolling 90-day window (Ahrefs Brand Radar Methodology). Otterly.AI takes the opposite approach — you manually add or discover the prompts you want tracked, and it runs them daily across ChatGPT, Google AI Overviews, Perplexity, and Microsoft Copilot, with Claude and Gemini available as add-ons, reporting mentions, citations, and sentiment weekly (Otterly.AI prompt monitoring docs). Semrush's Prompt Tracking runs the same target prompts daily against Google AI Mode, AI Overviews, Gemini, and ChatGPT Search, using the same crawling infrastructure as its traditional rank tracker, against a claimed database of 289 million prompts (Semrush: where AI Visibility Toolkit data comes from).

None of this is pulled from an official API most of these platforms expose for this purpose — it's automated querying of the consumer-facing product, parsed after the fact.

Comparing the common approaches

ApproachEngines coveredWhere prompts come fromCadenceWhat you actually get
Manual/DIY spreadsheetWhichever you run by handYou write themWhenever you remember toFull control, zero cost, doesn't scale past ~20 prompts
Otterly.AIChatGPT, Google AI Overviews, Perplexity, Copilot (Claude/Gemini as add-ons)You add or discover themDaily monitoring, weekly reportsPrompt-level control, narrower default engine set
Ahrefs Brand RadarGoogle AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini, CopilotAuto-generated from People Also Ask + Ahrefs' keyword corpusMonthly refresh, 90-day windowBroad, real-demand coverage; less control over exact wording
Semrush AI ToolkitGoogle AI Mode, AI Overviews, Gemini, ChatGPT SearchCustom prompts, position-tracking styleDailyIntegrates with existing SEO tracking; ChatGPT coverage limited to Search mode
ProfoundChatGPT, Google AI Mode (prompt/topic-based)Prompt and topic tracking you configureOngoingDeep citation-source taxonomy (8 categories); score is explicitly framed as directional

No two of these will give you the same "visibility score" for the same brand, because they're sampling different prompt sets on different schedules against engines that are themselves non-deterministic. That's not a bug in any one tool — it's the nature of measuring something with no ground-truth API.

The limitations nobody puts on the pricing page

There's no official measurement API, full stop. Even Google — which controls both the search engine and the AI feature — only gave publishers an AI-specific reporting surface in mid-2026, and it's a partial one. Search Console's Generative AI features report, which began rolling out in beta on June 3, 2026, shows impressions when a URL appears inside an AI Overview or AI Mode response, but explicitly has no query-level breakdown, no click data, and no CTR — you can see that you showed up, not what someone asked to trigger it (Am I Cited: what Google's new AI report still can't tell you). If the platform that owns the data won't expose query-level detail, no third-party vendor can either — they're all working from parsed, reconstructed responses.

Answers are noisier than they look. A 2026 statistical analysis of citation behavior across Perplexity, OpenAI's SearchGPT, and Google Gemini sampled the same topics daily over nine days and, in one arm of the study, at ten-minute intervals. It found citation distributions follow a power-law shape and that rankings of which domains get cited are unstable across repeated samples — not just at the margins, but throughout the frequently-cited set. The paper's conclusion is blunt: single-run visibility metrics, the kind almost every commercial tool reports from one batch of queries, "provide a misleadingly precise picture" and should be reported with confidence intervals, not point estimates (Sielinski, "Quantifying Uncertainty in AI Visibility").

Prompt-set choice moves the number as much as anything else. A separate 2026 study that analyzed over 100,000 prompt responses across 100+ tracked brands found visibility splits sharply by brand tier — roughly 73% mention rate for global brands, 44% for mid-market, 11% for niche brands — and, notably, that sentiment "flips" about 6.7 times more often between runs than whether the brand is mentioned at all (Kumar, "Generative Engine Optimization at Scale"). If your prompt library skews toward brand-name queries versus category queries, your mention rate will look completely different, independent of anything you actually did.

The academic baseline for "can you move the number" is more modest than marketing claims suggest. The 2024 Princeton/IIT Delhi paper that coined "Generative Engine Optimization" ran roughly 10,000 queries across nine benchmark datasets and found the strongest content-optimization techniques produced a 30–40% relative lift on their visibility metric — and that figure is a maximum from the best-performing methods under favorable conditions, not an average outcome you should expect (Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735). Later corrective work has pushed back further, finding several popular "GEO tactics" don't reliably help and that plain source relevance still does most of the work.

A practical way to run this without overclaiming precision

  1. Build a prompt set of 30–50 queries, split across direct brand questions, category/problem questions, and named comparisons against 2–3 real competitors. Pull some phrasing from actual search queries (Google Search Console, "People Also Ask," or your own sales call transcripts) rather than writing all of them from a marketing brief.
  2. Run every prompt at least twice per cycle, not once. Given the documented run-to-run variability, a single pass tells you almost nothing about whether a change in the number is signal or noise.
  3. Log four things per response: mentioned (yes/no), cited-with-link (yes/no, and to what URL), competitors also mentioned, and sentiment (positive/neutral/negative).
  4. Calculate mention rate and citation rate separately — they answer different questions. Mention rate tells you if the model knows who you are; citation rate tells you if it's actually pointing users to your site.
  5. Re-run on a fixed cadence (weekly or monthly, matched to how often you publish or change positioning) and track the trend line, not any single reading.
  6. Treat the absolute score as directional. Compare your own numbers over time and against named competitors in the same run — don't compare your score in one tool against someone else's score from a different tool, different prompt set, and different day.

The honest summary: nobody, including the platforms themselves, can currently give you an exact count of how often your brand appears across ChatGPT, Perplexity, Gemini, and Google's AI features. What you can get is a reasonably reliable trend, built from a consistent, repeated prompt set, that tells you whether visibility is moving up or down relative to competitors — which is usually the actionable question anyway.

nqzai's AI Visibility Score works within these same constraints rather than pretending they don't exist: it runs buyer-intent prompts against ChatGPT, Claude, Gemini, and Perplexity directly, plus Google AI Overviews as its own tracked surface, and combines entity recognition, page-readiness, and multi-engine citation checks into a composite 0–100 score with prioritized next actions — available through nqzai's AI search optimization. It's built to answer "is this moving and why" rather than to claim a precision the underlying data doesn't support.

FAQ

How do I check if ChatGPT mentions my brand?

Ask it directly with a handful of realistic prompts (your brand name, your product category, "vs" comparisons against competitors), and repeat the same prompts on a schedule — a single query gives you one noisy data point, not a reliable answer.

What's the difference between mention rate and citation rate?

Mention rate is how often your brand name appears anywhere in the answer. Citation rate is how often that mention comes with an actual link or named-source attribution back to your site. A brand can have a high mention rate and a low citation rate if the model is recalling it from training data rather than an active web source.

Is there a free tool to check AI search visibility?

Several vendors (Ahrefs, Semrush) offer limited free AI-visibility checkers that run a handful of prompts against your domain. They're useful for a one-time snapshot but won't show trends, since tracking over time requires a paid plan.

Why do different AI-visibility tools give me different numbers for the same brand?

They use different prompt sets, different sampling frequencies, and query the engines at different times — and answer-engine responses are documented to vary meaningfully between repeated runs of the identical query. Compare a tool's number against itself over time, not against a different tool's number.

How often should I re-measure AI search visibility?

Weekly if you're actively publishing or changing positioning; monthly is a reasonable floor. What matters more than frequency is running the same prompt set consistently so the trend is comparable.

Can Google Search Console show me which queries triggered my AI Overview appearance?

Not yet as of its Generative AI features report launched in mid-2026 — it shows impression counts by page but no query-level breakdown, so you can confirm you appeared but not what someone asked to trigger it.