AI Share of Voice: How to Calculate It, and Where It Misleads
Ask a language model who the best accountants in Leeds are. Write down the names. Ask again — same model, same words, same day — and the list can come back different. Not reordered. Different names.
That single observation is most of what “AI share of voice” is trying to measure, and most of why the number is harder to trust than it looks. It’s a real metric, tracked by real tools, with a real formula behind it. It’s also being reported to one decimal place for something that doesn’t hold still between two asks of the identical question.
What is AI share of voice supposed to measure?
The definition most vendors converge on: how often, and how prominently, a brand turns up in AI-generated answers relative to competitors, across a set of prompts someone in that category might actually type. HubSpot’s own glossary states it almost word for word — the share of a category’s AI answers a given brand is winning (HubSpot).
It’s borrowed directly from traditional media share of voice, where the concept has a fixed, auditable total: total advertising spend in a market, or total impressions across a set of channels. That total exists whether or not anyone measures it. AI share of voice borrows the name and the percentage format, but not that property — which is where the trouble starts.
How is it actually calculated?
Semrush’s guide gives the plain version: count how many times answer engines mention brands across a defined set of prompts, divide one brand’s mentions by the total, multiply by 100. A hundred mentions across a category, twenty-five of them for one brand, and that brand’s share is 25% (Semrush). Semrush’s own commercial tool layers a position weighting on top — being named first counts for more than being named fifth — but the base arithmetic is the same everywhere: mentions over total mentions.
That formula is straightforward. The two words doing the least work in it are “total mentions.” The broader mechanics of how a model reads and ranks a brand matter here too — share of voice is downstream of all of it, not a separate signal.
Why doesn’t the total mean what it looks like it means?
Because nobody is counting every AI answer ever given about a category. What actually gets counted is a specific, finite set of prompts, run against a specific set of models, on a specific day — and whichever competitor names happen to show up in that specific batch of answers. Change the prompts, the models, or the day, and the total changes with it, even if nothing about the actual market did.
One recommendation-style check built into a visibility-audit engine illustrates the scale of that finiteness directly. For a single category, it runs three differently-worded versions of a buyer’s question — “best X,” “who are the best X providers,” “I need an X, who would you recommend” — against three separate models, twice each. Eighteen answers, for one category, before it reports one number. A competitor’s name is only added to that number if it appears word-for-word inside one of those eighteen answers; nothing is inferred or guessed at. Eighteen is not a small sample by casual standards, and it is still short of the three-to-five-repeats-per-prompt floor llmpulse.ai names as the minimum for a position-weighted score to stop bouncing around (llmpulse.ai). Whatever total that run produces is real, and it is also just that run’s total — not the category’s.
Why does asking twice change the answer?
Because the answer is sampled from a probability distribution, not retrieved from a fixed record. llmpulse.ai’s own explanation of this is direct: run the identical prompt twice and the brand order frequently changes, because that’s how the underlying model generates text at all (llmpulse.ai). The recommendation-style check above is deliberately run at a temperature setting above zero for this reason — turned up on purpose, not left as a default, specifically so a report doesn’t mistake one greedy, most-likely answer for a settled fact about what a model “thinks.” A single answer to an open-ended category question is a sample of one. A related look at AI visibility scoring covers why that same instability runs through more than just share of voice — this is the sharper version of that same problem, applied to a number vendors present as a competitive ranking.
Does a higher number mean more customers?
Not by itself. Share of voice counts whether a model volunteers a brand’s name unprompted — a recommendation signal, closer to being mentioned by a well-informed friend than to being clicked on. It says nothing about whether the person asking then visited the site, called the number, or bought anything. A brand can lead its category’s share of voice and still see no measurable change in leads, and a brand with a modest share can still win the customers who found it through a channel none of this counts at all. The percentage answers “how often does a model say my name,” not “how often does that turn into a customer” — and the second question needs its own evidence, not an inference from the first. A free scan shows the deterministic half of that picture — crawler access, structured data — for nothing; the recommendation sampling behind a share-of-voice figure is a deeper, paid layer on top of it, not a substitute for it.
None of this makes the metric worthless. It makes it a directional read from a small, dated, prompt-specific sample — closer to a poll than a census — and worth treating that way: a range, checked periodically, across more than one phrasing and more than one model, rather than a single percentage compared to the decimal point against last month’s.
Frequently asked questions
What is AI share of voice?
A metric that measures how often a brand appears in answers from tools like ChatGPT, Gemini, and Perplexity, relative to competitors, across a defined set of category prompts. HubSpot's glossary and Semrush's own guide both define it this way — the percentage of AI-generated responses that mention a brand out of all brand mentions in that same set.
How is AI share of voice calculated?
The basic version is mentions divided by total mentions, times 100: if answer engines name brands 100 times across a set of prompts and one brand accounts for 25 of them, its share of voice is 25%. Semrush's own tool adds a position weighting on top of that base formula — how high a brand appears, not just whether it appears — but the underlying total is still whatever set of prompts and competitors that particular run happened to surface.
Why does the same prompt return different brands each time it's asked?
Because a language model's answer to an open-ended question is sampled, not looked up. llmpulse.ai's own guide to the metric puts it plainly: run the same prompt twice and the brand order frequently changes. Tools built to measure this ask each prompt more than once for exactly that reason — a single answer is one sample of a variable process, not a settled fact.
Does asking more times fix the number, or just describe it more precisely?
It describes it more precisely, up to a point — it doesn't remove the deeper problem. llmpulse.ai recommends at least three to five repeats per prompt before a position-weighted score stabilises. More repeats narrow the range around a percentage; they don't turn the total mentions being measured against into a fixed, comparable universe.
Does a higher AI share of voice mean more customers?
Not directly. The metric counts how often a model volunteers a brand's name when asked a category question that never names it — a recommendation signal, not a transaction. A brand can be named in every sampled answer and still convert nobody, or be named rarely and still win the customers who found it another way. It's one input, not a proxy for revenue.
