TL;DR
Running the same AI prompt twice can produce different citations because even at temperature zero, production LLM APIs are not bit-for-bit reproducible due to batching and inference nondeterminism, and Google’s AI Overviews use a “query fan-out” technique that issues multiple related searches, making results vary between runs minutes apart.
The most common failure mode in AI-visibility reporting is someone writing ten prompts they suspect will surface their own brand, running each once on a single logged-in account, and claiming “cited in 8 of 10 answers” — that measures the prompt-writer’s intuition, not real customer query space. OpenAI’s own documentation
What an AI citation research method is
Direct answer: An AI citation research method is the documented procedure you use to choose which prompts get run against AI answer engines — ChatGPT, Google's AI Overviews and AI Mode, Perplexity, Gemini, Copilot — in order to measure whether and how often a brand, product, or page gets cited in the response. The method is not the citation count. It's the sampling design that produced the prompt list, run on a defined cadence, against a defined set of engines, with the personalization and randomness controlled for as much as each platform allows.
That distinction matters because the common failure mode is specific and repeatable: someone writes ten prompts they already suspect will surface their own brand, runs each one once, on one logged-in account, and reports "we're cited in 8 of 10 AI answers." That number describes the prompt-writer's intuition about their own product, not the brand's visibility across the query space a real customer would actually type. It's a screenshot with a spreadsheet attached, not research.
Why the prompt set is the entire experiment
In survey research, the sample is the thing that determines whether a finding generalizes to anything beyond the people you asked. Pew Research Center's methodology documentation frames this as "total survey error" — coverage error, sampling error, and selection bias all compound before a single answer is even analyzed, and selection bias specifically occurs "when the kinds of people who choose to participate are systematically different from those who do not" (Pew Research Center, "Evaluating Online Nonprobability Surveys," 2016). Swap "people who choose to participate" for "prompts someone chose to write" and you have the exact failure mode of most AI-visibility reporting: the sample is opt-in, self-selected by the person with the incentive to see a good result, and none of that gets disclosed.
The same problem shows up in academic LLM evaluation. A 2024 paper on benchmark robustness notes that "models leverage spurious statistical relationships in the benchmark datasets, leading to overestimated performance," and that model results are "highly sensitive to minor changes in benchmarks" (Examining the Robustness of LLM Evaluation to the Distributional Assumptions of Benchmarks, arXiv:2404.16966). If academic benchmarks built by researchers with statistics training are still fighting sampling artifacts, a ten-prompt list a marketer wrote in twenty minutes is not a reliable instrument.
Four forces that move the same query under you
Direct answer: Before designing a sample, it helps to understand why running the identical prompt twice can produce two different answers with two different citation sets.
Personalization. Google has been rolling out "Personal Intelligence" across AI Mode and Gemini, which connects a user's Search history, Gmail, and Photos to tailor responses — Google's own product blog describes it as letting users "connect the dots across Google apps... to provide responses that are uniquely relevant to them" (Google, "Personal Intelligence expands in the U.S.," blog.google). OpenAI made a parallel move in April 2025 when it updated ChatGPT's memory to draw on the full history of past conversations rather than only explicitly saved facts, so two accounts asking an identical question can be nudged toward different phrasing and different sources based on what each account has discussed before.
Sampling randomness, even at temperature zero. Thinking Machines Lab's engineering team wrote that "it's remarkably difficult to get reproducible results out of large language models," and that even at temperature 0 — theoretically the deterministic, greedy-decoding setting — production LLM APIs are not bit-for-bit reproducible, largely because of how requests are batched together at inference time (Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference," 2025). A follow-up study that probed token-level probabilities rather than just final text confirmed the effect is real and can meaningfully change output when a token's probability sits in the 0.2-0.8 range rather than near a hard 0 or 1 (Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism, arXiv:2601.06118). OpenAI's own documentation is candid about the ceiling here: even with a fixed seed parameter and identical settings, the API can only be "mostly deterministic," not deterministic (OpenAI Cookbook, "How to make your completions outputs consistent with the new seed parameter").
Live retrieval and query fan-out. Google's Search Central documentation states that "both AI Overviews and AI Mode may use a 'query fan-out' technique — issuing multiple related searches across subtopics and data sources — to develop a response," and explicitly warns that AI Overviews and AI Mode "may use different models and techniques, so the set of responses and links they show will vary" (Google Search Central, "AI Features and Your Website"). Because the underlying search happens live, the exact set of pages an answer engine pulls from can shift between two runs of the same query minutes apart, independent of anything about the user.
Silent model and index updates. Providers routinely swap model weights or serving infrastructure behind an unchanged API name. OpenAI exposes a system_fingerprint value specifically so developers can detect when "OpenAI has made changes on their end" that could alter output for identical requests — an implicit admission that the same model name doesn't guarantee the same behavior over time.
Any one of these four forces alone would justify sampling instead of a single spot-check. Together, they mean a one-time run of a handful of prompts measures a specific account, on a specific day, hitting a specific model snapshot, retrieving a specific live index — a sample size of one, dressed up as a finding.
Comparison of prompt-sampling approaches
| Sampling approach | How prompts get chosen | Primary bias introduced | When it's defensible |
|---|---|---|---|
| Vanity / branded-only | Prompts written to include the brand or a close variant | Confirms what you already believe; ignores how buyers actually search | Never as a standalone visibility claim — only as one labeled slice of a larger set |
| Convenience sample | Whatever prompts the team happens to think of in a meeting | Reflects internal jargon and internal mental models, not customer language | Early brainstorming only, not reporting |
| Competitor-mirrored | Copy prompts a competitor or vendor published in a case study | Inherits someone else's category framing and blind spots | Useful as a cross-check, not a primary source |
| Keyword-to-prompt conversion | Take existing SEO keyword list and reword as questions | Skews toward transactional, high-volume terms; misses conversational and troubleshooting phrasing unique to chat interfaces | Good supplementary layer, weak as the sole source |
| Stratified, intent-based sample | Deliberately built across awareness, comparison, transactional, and troubleshooting intents, branded and unbranded tracked separately | Requires more upfront design work; still limited by which categories you chose to stratify on | The defensible default for ongoing tracking |
| Logged-out / non-personalized run | Same stratified set, executed without account history or location signal where the platform allows it | Understates real-world personalization effects; some platforms won't allow fully logged-out AI answers | Best available proxy baseline — pairs with logged-in spot checks, doesn't replace them |
Step-by-step: building a defensible sample-prompt set
- Define the query space before writing a single prompt. List the topics, product categories, and buyer questions the brand should plausibly be cited for, independent of whether you think you'd win. Treat this as a coverage map, not a wishlist.
- Separate branded from unbranded prompts, and never blend the two in one headline number. A prompt containing your brand name measures something different — recall and disambiguation — from a prompt that never mentions you and tests whether the model surfaces you unprompted.
- Source prompt language from real customers, not internal vocabulary. Support tickets, sales call transcripts, review site questions, and community forum threads produce phrasing that matches how people actually ask AI systems questions, which is often more conversational and less keyword-shaped than SEO-derived phrasing.
- Stratify by intent stage. At minimum: awareness ("what is X"), comparison ("X vs Y"), transactional ("best X for Z"), and troubleshooting/support. This is the same logic Google's own query fan-out documentation describes AI systems using internally to decompose a single query into subtopics — your sample should mirror that structural diversity rather than testing one intent nine different ways.
- Fix the sample size and stick to it before you see results. Decide up front how many prompts per stratum you're running — not "however many needed until the numbers look good."
- Run each prompt more than once, across sessions. Given the documented non-determinism in LLM sampling, a single run per prompt cannot distinguish signal from noise. Running the same prompt multiple times and reporting a citation rate, not a citation flag, is the minimum bar the nondeterminism research supports.
- Run a logged-out or non-personalized baseline alongside logged-in spot checks. Rank-tracking vendors that monitor AI Overviews have documented that personalized AI answers are, by design, invisible to any tool that can't authenticate as a specific real user with real history — Advanced Web Ranking's own study notes that AI Overviews "only appear for logged-in users and do not show up in incognito searches," which caps what any automated tracker can see (Advanced Web Ranking, "AI Overview Study for 8,000 Keywords in Google Search"). Treat the non-personalized run as your comparable baseline, and treat any logged-in spot check as illustrative, not statistical.
- Repeat on a fixed cadence, not once. Because retrieval is live and models update silently, a prompt set is a snapshot the moment you run it. Recurrence — weekly or monthly, held constant — is what turns a snapshot into a trend.
- Version and publish the prompt list itself. Date-stamp the exact prompt wording, the engines tested, and any exclusions. If the list can't be handed to a skeptical colleague to reproduce, it isn't a method — it's an anecdote with a chart.
What this doesn't guarantee
Direct answer: A disciplined sampling method reduces bias; it doesn't eliminate the underlying volatility. Be explicit about the limits:
- It cannot promise a stable, single "visibility score." Because personalization, live retrieval, and model updates are all real and documented, the honest output of any AI citation study is a rate with a confidence range and a date, not a fixed number that holds until you check again.
- It cannot see what a logged-in, personalized user with real history actually gets. Any method that runs from a clean account is measuring a proxy for the personalized experience, not the personalized experience itself — a gap the industry's own rank-tracking research openly acknowledges rather than a solvable engineering problem.
- It cannot guarantee that citation optimization tactics move the number. The NeurIPS 2025 C-SEO Bench study tested conversational-SEO tactics directly and found most "are largely ineffective," concluding that "the initial ranking of documents retrieved by the search engine plays a far more dominant role" than any content-side trick (C-SEO Bench: Does Conversational SEO Work?, arXiv:2506.11097). A good sampling method tells you honestly whether visibility changed — it doesn't manufacture the change.
- It cannot substitute for a large enough sample on a narrow question. The original Princeton GEO study needed roughly 10,000 queries to produce statistically robust conclusions about which content strategies helped (GEO: Generative Engine Optimization, arXiv:2311.09735); a 40-prompt brand-tracking set is proportionate for ongoing monitoring, not for drawing firm causal conclusions about what worked.
- It goes stale. A prompt list built around today's product line, competitors, and customer language needs to be revisited as the market and the models change — a method with no version history is already an outdated one.
Where nqzai fits
Direct answer: nqzai's AI-visibility tooling is built to run a structured, stratified prompt set — branded and unbranded, split by intent stage — against multiple AI answer engines on a repeating schedule, then track citation presence and share of voice over time rather than producing a single flattering screenshot. It surfaces the underlying prompt list and run dates alongside the results, so a citation-rate claim can be checked against the sample that produced it instead of taken on faith.
FAQ
How many prompts do I need for a defensible sample?
There's no universal minimum, but the number should be set by your stratification design, not by how many prompts produce a good-looking result. A reasonable starting point for ongoing brand monitoring is enough prompts per intent stage (awareness, comparison, transactional, troubleshooting) to detect a meaningful shift — typically dozens per stratum rather than a handful total — repeated across multiple runs, not run once.
Should branded and unbranded prompts be reported together?
No. They measure different things. A prompt with your brand name in it tests recall; a prompt without it tests unprompted discovery. Blending them into one citation-rate number hides which one is actually driving the result.
Why does the same prompt give different citations when I run it twice?
Documented sampling randomness in LLM inference (even at low temperature), account-level personalization from search and conversation history, live retrieval pulling a different set of pages at different moments, and silent backend model updates all contribute — see the sources on nondeterminism and personalization above. This is expected behavior, not a bug in your testing.
Is a logged-out prompt run a fair test?
It's a fair baseline, not a complete picture. Logged-out runs strip out personalization, which makes them reproducible and comparable over time — but they can't see what a real, signed-in user with search and browsing history actually receives, since platforms like Google's AI Mode explicitly personalize based on that history.
Do AI citation-boosting tactics actually work?
The evidence is mixed and tactic-dependent. Princeton's original GEO research found some content changes (adding citations, statistics, and quotations) improved visibility in a controlled benchmark by roughly 30-40%. A later, larger NeurIPS study testing tactics head-to-head against real search engines found most had little to no effect once retrieval ranking was accounted for. Treat any single tactic's claimed lift skeptically until you've tested it against your own prompt set.
How often should I re-run the prompt set?
On a fixed cadence — weekly or monthly is common — held constant over time so a change in the numbers reflects a change in visibility rather than a change in when you happened to check. Ad hoc, whenever-you-remember checks make trend analysis impossible.



