Does ChatGPT Give the Same Answer to Everyone?
In November and December 2025, Rand Fishkin of SparkToro and Patrick O’Donnell of Gumshoe.ai ran a study that put 2,961 prompts through ChatGPT, Claude and Google’s AI Overviews, repeating each one 60 to 100 times. The prompts asked for recommendations: best chef’s knives, best divorce lawyers, best noise-cancelling headphones. Two runs of the exact same question returned the exact same list of brands in under 1% of tries. The same list in the same order came back in under 0.1% — less than one time in a thousand.
That is the answer to the question in the title. ChatGPT, asked twice, is not answering twice from the same script.
Why the answer moves between two identical questions
Most consumer AI products don’t generate text by always picking the single most probable next word. They sample — choosing each word as a weighted guess among several plausible options, with a setting (temperature) that controls how much latitude that guess gets. A setting of zero would make the model nearly deterministic, repeating the same answer almost every time. Consumer chat products don’t run at zero. They run with enough latitude that the wording drifts, and sometimes the actual list of names drifts with it.
This isn’t a bug being patched out. It’s closer to a coin that’s deliberately a little bent, on purpose, because a model that always gave the single most statistically likely phrasing would sound stilted and repetitive over millions of conversations. The same mechanism that keeps a chatbot’s prose varied is the one that makes its brand-recommendation lists wobble.
One answer is a sample, not a verdict
A business that types its own name into ChatGPT once, sees itself mentioned, and concludes it has cracked AI search has learned almost nothing. It’s the equivalent of asking one customer and generalizing to a market. The SparkToro figures put a number on exactly how little: if the same question asked twice produces the same list under 1% of the time, then one good answer today says very little about what tomorrow’s identical question will return.
The only honest way to measure this is to ask repeatedly and report a rate — recognized in 6 answers out of 10, say, rather than recognized or not. Measuring AI visibility this way behaves less like a fact and more like a batting average, and the same logic is what turns a raw mention count into something closer to share of voice: a proportion out of many tries, not a single yes or no.
How a measurement actually handles the noise
A check that tries to measure this doesn’t ask once. It puts the same question to more than one model — not just ChatGPT, but a small roster of different systems — and asks each model more than once, with sampling deliberately left turned up rather than flattened to zero, because flattening it would hide the exact variance being measured. Each answer that comes back gets sorted into one of three outcomes: it names the business outright, it answers with a generic definition of what the name usually means (a sign the model doesn’t actually know the business and is just parsing the words), or it says plainly that it has no information. Only the first counts as real recognition.
That still leaves a limitation worth stating rather than hiding. Two answers per model is enough to notice that an answer moved between tries. It is not enough to pin down a tight, confident number for how often it moves — a rate measured from two samples carries real uncertainty of its own, and a business reading “recognized in 1 of 2 tries” should treat that as a rough signal, not a precise score. A tighter number needs more repeat questions than most audits, this kind included, currently run, because every repeat is itself a separate billed request to a model provider. This part of a check — a direct, un-primed question put straight to a model with no supporting material attached — also only runs in a deeper, paid pass. A free scan skips it entirely and sticks to what a crawler can read from the site in a single visit, which is a different, narrower kind of evidence than asking a model what it already believes.
What a business can actually act on
The randomness itself isn’t something a business gets to switch off. It’s a deliberate property of how these systems generate text, not a setting hidden in some dashboard waiting to be found. What moves is the rate — how often, across repeated tries, a model lands on recognizing a specific business rather than reciting a generic definition of a common phrase or admitting it has nothing. A business whose name, services and location are stated clearly and consistently across its own site gives every one of those repeated tries slightly better odds of landing on the right answer. None of that guarantees any single try. It shifts the average over many of them, which is the only number that was ever going to mean anything here. The broader mechanics of how a model reads a page before it ever answers a question about a brand sit underneath that average, whether the question gets asked once or a hundred times.
A screenshot of ChatGPT naming — or skipping — a business proves less than it looks like it proves. The number worth trusting is the one built from asking the same question several times over, and even that number arrives with a margin of error, because knowing an answer moved is not the same as knowing by how much.
Frequently asked questions
Does ChatGPT give everyone the same answer?
No. A study that ran 2,961 identical prompts through ChatGPT, Claude and Google's AI Overviews 60 to 100 times each found the same brand-recommendation list came back in under 1% of repeat tries, and the exact same order in under 0.1%. Two people typing the same question rarely see the same list of names.
Why does ChatGPT answer differently each time, even for the same question?
Most consumer chat products generate text by sampling, choosing each next word as a weighted guess rather than always picking the single most likely one. That sampling is deliberately turned up above zero, so the exact wording, and sometimes the exact list of names, can shift between two runs of the identical prompt.
If answers vary, how can a business measure whether it gets mentioned by AI at all?
By asking the same question repeatedly and reporting a rate rather than a single pass/fail. One answer is a sample, not a verdict. A measurement built this way asks the same question of more than one model, more than once each, and reports how often a brand's name showed up against how often it didn't, rather than trusting any single try.
Does this variance also happen with Claude and Google's AI Overviews?
Yes. The same study found it across all three systems it tested, not just ChatGPT. Which brand reads as the top answer can change depending on which engine was asked, on top of the variance any one engine shows against itself.
Can a business do anything about the randomness itself?
Not the randomness. Sampling is a deliberate design choice in these systems, and a business can't turn it off from the outside. What's open to influence is the rate: a brand with clear, consistent, well-structured facts about what it does and where tends to get named more often across repeated tries, even though no single try is ever guaranteed.
