TL;DR

ChatGPT gives different answers to the same question about brands across repeated runs, with research showing it flips on factual yes/no questions only 73% of the time and brand mentions can vanish entirely on a second try. A single "I checked ChatGPT and we weren't mentioned" report is worthless because it's built on a sample size of one, measuring noise as signal. To get reliable data, run 15–25 specific buyer prompts 5–10 times each, logging not just whether your brand appeared but its position, sentiment, competitors named, and whether a URL was cited.

The bottom line: treat ChatGPT brand visibility as a distribution, not a ranking, and use the patterns to diagnose whether you're actually invisible or just unlucky in one run.

Ask ChatGPT the same question about your category twice in a row and you can get two different answers — different brands, different order, sometimes a mention that vanishes entirely on the second try. That's not a bug you can wait out. It's how the system works, and it means most "I checked ChatGPT and we weren't mentioned" reports are built on a sample size of one. This is a method for doing better: a fixed prompt panel, a consistent way to run it, a way to log what actually happened (not just yes/no), and a framework for figuring out why a brand is missing before you conclude it's actually invisible.

Why one prompt is not an audit

ChatGPT's outputs are probabilistic by design. The model samples from a distribution over possible next tokens rather than computing one fixed answer, and that sampling process is the source of variation you see run to run. Counterintuitively, this doesn't fully disappear even at settings that are supposed to be deterministic — engineering research from Thinking Machines Lab found that even greedy, temperature-zero decoding on the same model can still produce different outputs across runs, likely due to floating-point non-associativity and how concurrent GPU kernels execute in whatever order they finish ('Defeating Nondeterminism in LLM Inference'). A companion academic study measuring "deterministic" LLM settings in practice found the same pattern: nominally deterministic configurations still drift across repeated calls (arXiv:2408.04667).

This isn't just an infrastructure curiosity — it shows up at the level of what gets said. Washington State University researchers asked ChatGPT the same yes/no research question ten times per hypothesis, across more than 700 hypotheses, and found the model gave consistent answers only about 73% of the time — flipping between "true" and "false" on identical prompts within the same batch of ten runs (WSU: "AI gets a D"). Industry analysis of repeated brand-visibility prompts finds the same instability in mention behavior specifically, not just factual answers (Search Engine Land: "What repeated ChatGPT runs reveal about brand visibility"). If a factual claim can flip five times in ten runs, a brand mention — which depends on far more marginal signal — can certainly flip too. Treating one run as your "ranking" is measuring noise and reporting it as signal.

How ChatGPT decides what to say about a brand

Before building a panel, it helps to know what's actually happening under the hood. In its default mode, ChatGPT answers from parametric knowledge learned during training — no retrieval at all. When search is triggered (automatically, or because the user enables it), ChatGPT rewrites the request into one or more targeted queries, sends them to search partners, retrieves results, and grounds its answer in what comes back, with source links available in a sidebar (OpenAI: "Introducing ChatGPT search"; OpenAI Help Center). Reporting on the underlying mechanics describes this as a "fan-out" process — an initial query followed by narrower follow-up queries as the model refines what it's looking for (Search Engine Land: "Inside ChatGPT Search").

This matters for your audit design in a concrete way: whether search is on, whether it's triggered at all, and which follow-up queries the model chooses to run are themselves variable and partly outside your control. A prompt that reads identically to a human can trigger retrieval in one run and answer from memory in another. That's one more reason a single run tells you almost nothing — you don't even know reliably whether it was answered from trained knowledge or from a live search.

Step 1: Build a fixed prompt panel

Direct answer: Write down 15–25 prompts a real buyer would type, and don't touch the wording once you start. A workable panel spans four categories:

CategoryPurposeExample shape
Category/discoveryTests unaided visibility"What are the best tools for [job-to-be-done] for [company type]?"
ComparisonTests whether you appear against named competitors"How does [Competitor] compare to alternatives for [use case]?"
Problem-solutionTests visibility when the buyer hasn't named a category yet"We're struggling with [specific pain point] — what should we look into?"
Branded / evidence-seekingTests what ChatGPT knows and cites about you directly"What is [Your Brand] known for?" / "What sources describe [Your Brand]?"

Weight the panel toward category and problem-solution prompts. Branded prompts are diagnostic — they tell you what ChatGPT already associates with your name — but if that's all you track, you'll only ever learn about people who already know you exist. Keep prompts specific ("expense automation software for 200-person distributed teams," not "expense software") since specificity is closer to how a real buyer phrases things and less likely to collapse into generic top-of-category answers.

Step 2: Run it consistently, across conditions

Direct answer: For every prompt, run it multiple times — five to ten repetitions is a reasonable floor given how much run-to-run variance the research above documents. Where possible, vary the conditions deliberately rather than letting them vary by accident: logged out vs. logged in, search off vs. search on if you can control it, and if you serve multiple regions, from more than one locale. Record the date, the exact prompt text, and the full response verbatim — not a paraphrase — so you can re-check your own coding later.

Step 3: Log patterns, not a yes/no

A mention/no-mention checkbox throws away almost everything useful. For each run, capture:

FieldWhy it matters
Mentioned (Y/N)The baseline signal
Position in the list/answerEarly mentions read as stronger endorsement than a tail entry
Framing/sentimentNeutral listing vs. described as a leader vs. described with a caveat
Competitors named alongside youShows the comparison set ChatGPT considers you part of
Whether a URL was citedA citation is a stronger signal than an unsourced mention
Which source(s) were citedTells you why ChatGPT said what it said

Treat mentions, citations, and cited sources as three distinct things you're measuring, not one. A brand can be named with no source behind it (recalled from training), or cited from a single unflattering third-party page, or absent from the answer but present in the "sources" sidebar. Each of those is a different problem with a different fix.

Step 4: When you're absent, look for the evidence gap

Absence is a symptom, not a diagnosis. The foundational research on what makes generative engines cite a source — a controlled study out of Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi that benchmarked nine content interventions against 10,000 queries — found that classical SEO signals like keyword density barely moved citation likelihood, while content carrying statistics, direct quotations, and citations from credible sources produced measurable lift, in some cases over 30–40% (Aggarwal et al., "GEO: Generative Engine Optimization," KDD 2024). If your own site is the only place making your claims, that's a gap worth closing — the model has no independent evidence to draw on.

Independent, large-scale correlational data points the same direction from the other end. Ahrefs' analysis of 75,000 brands found that how often a brand is mentioned across the open web correlates with AI visibility far more strongly (0.664) than backlink counts do (0.218), and that brands in the top quartile for web mentions received roughly ten times more AI-answer mentions than the next tier (Ahrefs: "An Analysis of AI Overview Brand Visibility Factors"). A follow-up comparison across engines found this pattern holds for Google's AI Overviews specifically, with a notably weaker relationship for ChatGPT and Perplexity individually — a reminder that "visible in one AI surface" doesn't transfer cleanly to another (Ahrefs: "Google Seems More Biased Towards Big Brands Than ChatGPT and Perplexity"). Correlation isn't causation — brands that earn a lot of independent coverage also tend to have other things going for them — but it's a reasonable starting checklist. When a prompt panel shows consistent absence, check for: third-party coverage of your brand (reviews, comparisons, analyst mentions, forum discussion) outside your own site; whether competitor names appear in text alongside credible, citable claims where yours don't; and whether your own pages contain the kind of concrete, sourceable evidence — data, named case outcomes, direct quotes — that the GEO research found actually moves citation behavior.

Step 5: Report a mention rate, not a rank

Once you have logged results across runs and conditions, the number that means something is an aggregate: mention rate (the percentage of runs across the panel where you appeared), citation rate (the percentage where a source link accompanied the mention), and framing consistency (how often the tone stayed positive or neutral vs. drifted negative). None of these should be reported from a single run. If you only have budget for one thing, report mention rate as a percentage with the run count next to it — "mentioned in 6 of 10 runs across this panel" is honest; "we rank #2 in ChatGPT" from one screenshot is not, and it will not replicate the next time someone checks.

Cadence

Direct answer: Re-run the full panel monthly at minimum, since both the underlying model and the live web content it draws on change continuously. Keep the prompt wording frozen between runs so movement in your numbers reflects a real change in visibility rather than a change in how the question was asked. Revisit the prompt list itself on a quarterly basis to keep it aligned with how buyers actually phrase things — but do that as a deliberate new baseline, not a mid-quarter edit.

Sources: