AI-SEO & GEO

AI Share of Voice: A Repeatable Method, and the Sample Size That Actually Means Something

· · 12 min read

AI share of voice is the metric everyone reports and almost nobody can defend. Ask how the number was produced and the answer is usually a tool, a prompt list nobody can see, and a single run. That is not a measurement, it is a screenshot with a percentage on it. The engines answer differently every time, the brand set is whatever the vendor detected, and the sample is too small for the one-point changes that get presented as movement.

This article is a method you can run yourself and defend in a meeting: how to define the metric so it has a denominator, how to design the prompt sample, why repeats and confidence intervals are not optional, and how many responses you need before a month-over-month change means anything. To make the point concrete I ran the method on a live sample this week, 32 responses across ChatGPT and Gemini, and the variance in it is the argument. If you want the manual prompt matrix that this method formalises, it is in checking whether your business appears in ChatGPT and Perplexity; if you want the tools that automate collection, they are compared in the AI visibility software guide. This is the layer in between: what to measure, and how much.

Key takeaways

  • AI share of voice only means something once the brand set, the prompt set, the engines, the location and the repeat count are fixed and written down; change any one and the number is not comparable.
  • In a 32-response sample this week, the same prompt on the same engine an hour later named a different set of brands in 9 of 16 pairs, and one ChatGPT prompt went from naming four tracked brands to naming none.
  • With 16 responses per engine, the 95% interval on a 50% mention rate is 44 points wide; getting it under 10 points needs about 400 responses per engine.
  • Detecting a 10-point month-over-month change at conventional confidence needs roughly 400 responses per engine per month; below that, reported changes are mostly sampling noise.
  • Cited domains are far more volatile than brand mentions: 79 of the 114 domains cited in the sample appeared exactly once.

What AI share of voice measures, and why most numbers are unfalsifiable

Share of voice is a ratio. The numerator is how often your brand appears; the denominator is how often any brand in a defined set appears, across a defined set of prompts. Every word in that sentence is a choice, and the metric is only as reproducible as the choices are explicit.

Most published AI share of voice numbers fail on the denominator. If the set of competitors is whatever the tool’s entity detection found in the answers, the denominator changes every run and the ratio is meaningless. If the prompts are proprietary, nobody can test whether they represent what buyers ask. If each prompt was run once, the number carries no information about how stable it is. And if the vendor rolls mentions, citations, position and sentiment into a composite “visibility score”, the components cannot be audited separately, and the score cannot be compared with anyone else’s.

The fix is not a better tool. It is a protocol: the same brand set, the same prompt set, the same engines and locations, run the same number of times, scored the same way, with the raw responses kept. Tools can execute a protocol. They cannot substitute for one.

Definitions: three metrics, one unit of analysis

Fix the unit of analysis first. The unit is one response: one prompt, on one engine, from one location, at one time. Everything is a rate over responses, never a count of times a brand’s name occurred, because an engine that repeats a brand five times in one answer has not made it five times as visible.

  • Mention rate for a brand is the share of responses in which the brand is named at all. It is the most robust metric because it does not depend on what other brands did.
  • Share of voice for a brand is its mentions divided by the total mentions of all brands in the set, where each brand counts at most once per response. It is the competitive metric, and it moves when competitors move even if your own mention rate does not, which is why it should never be reported without the mention rates beside it.
  • Citation share is the share of responses in which a page on your domain appears in the cited sources. It measures something different from being named, and the two diverge constantly, so keep them apart.

Add position and recommendation strength as secondary scores if you need them, but score them from the raw responses, keep the raw scores, and resist the composite.

The design: what has to be fixed before the first prompt

The brand set. Choose it in advance from your competitive frame, not from the answers. Six to ten brands is enough for a category. Write down aliases, because “Ahrefs Brand Radar” and “Ahrefs” are one brand, and decide whether a brand named only as the subject of a comparison prompt counts as a mention. In my sample it does, and the rule is written into the scoring notes.

The prompt set. Sample from the prompt types buyers actually use: category discovery, comparison, alternatives, use case, budget, and how-to. Do not include your own brand in the prompt unless you are deliberately measuring branded demand, because a prompt that names you measures whether the engine can talk about you, not whether it chooses you. Twenty prompts is a pilot. Fifty is a baseline. The arithmetic below says how many you need for the precision you want.

Engines and locations. Each engine is its own population. Report per engine, never pooled, and pin the location, because answers differ by country. In this sample every response was requested from the United Kingdom.

Repeats and timing. Run every prompt at least twice per period, separated by at least an hour, and treat the two runs as the minimum estimate of within-period variance. A single run cannot distinguish a change from noise.

Scoring rules. Write them before scoring. What counts as a mention, how aliases resolve, whether a source description counts as a citation if the brand is not in the answer text. Then score from the response text, not from the tool’s summary.

What 32 live responses showed

The sample: eight prompts in the AI visibility tools category, run on ChatGPT and Gemini from the UK, twice each on 11 September 2026 with the second run starting about an hour after the first. Brand set: Profound, Peec AI, Otterly, Semrush, Ahrefs and HubSpot. Perplexity was in the design and is not in the data, because the collection API returned an error on all six attempts; that is itself a finding about running this at scale, and it is why the protocol has to record engine availability per run.

Mention rates across all 32 responses, with 95% Wilson intervals:

BrandMention rate95% intervalShare of voice
Semrush72%55% to 84%22.5%
Profound62%45% to 77%19.6%
Ahrefs59%42% to 74%18.6%
Peec AI56%39% to 72%17.6%
Otterly50%34% to 66%15.7%
HubSpot19%9% to 35%5.9%

Read the intervals before the ranks. Semrush leads Otterly by 22 points in mention rate, and the intervals still overlap. Only HubSpot is separable from the rest. A vendor dashboard would show this as a ranked bar chart; the honest version is one cluster of five brands and one outlier.

Per engine, the picture differs in ways the pooled number hides. On ChatGPT the leaders were Semrush at 69% and Ahrefs at 62%; on Gemini, Semrush at 75% and Profound at 69%. HubSpot was named in one ChatGPT response and five Gemini responses, because Gemini kept surfacing its free grader. Five of the 32 responses named no tracked brand at all, and all five were ChatGPT: two prompts about method and budget produced answers built entirely around brands outside the set. That is a second reason to fix the set in advance. Across the sample, 54 brands outside the set were named, and a detected-brand denominator would have absorbed every one of them.

Run-to-run agreement is the number nobody publishes. For each prompt and engine I compared the set of tracked brands named in run one with run two. The mean Jaccard overlap was 0.76, seven of the 16 pairs were identical, and nine were not. The largest flip was prompt five on ChatGPT, a request for an agency platform, which named Profound, Otterly, Semrush and Ahrefs in the first run and none of them an hour later, recommending five different products instead. Ahrefs flipped in five of the 16 pairs, more than any other brand. Nothing changed on any of those brands’ websites in that hour. That is the noise floor a monthly report has to clear.

Citations were more volatile still. The 32 responses cited 114 distinct domains across 171 citation events, and 79 of those domains were cited exactly once. Comparing the cited domains for the same prompt and engine between runs gave a mean overlap of 0.31, against 0.76 for brand mentions. Eleven domains were cited by both engines. Citation share is worth tracking, but at this sample size it is a list of examples, not a rate.

Sample size: the arithmetic

A mention rate is a proportion, so its uncertainty follows from the number of responses. The table gives the approximate width of the 95% Wilson interval for a brand at a 50% mention rate, which is the worst case, at various sample sizes per engine.

Responses per engineInterval width at 50%
1644 points
5027 points
10019 points
20014 points
40010 points

Two consequences follow. First, a baseline that claims a brand “has 38% share of voice” from 20 prompts run once is claiming a precision it does not have; the interval on that number is wider than the gap between most brands in any category. Second, change detection needs more than the baseline. Comparing two monthly samples of 100 responses per engine, the difference in mention rate has to be about 20 points before it is distinguishable from sampling noise at conventional confidence. At 400 per engine per month it is about 10 points. Below that scale, the small monthly movements that get reported as wins and losses are mostly the engines being non-deterministic.

Because responses to the same prompt are not independent, count prompts before repeats. Fifty prompts run eight times is not the same as 400 prompts run once; it is 50 clustered observations with a variance estimate. Grow the prompt set first, keep at least two repeats for the variance, and compute intervals by bootstrapping over prompts rather than over responses. The same bootstrap logic applies to AI Overview CTR analysis, where the query is the cluster.

Scoring beyond the mention

Once the sample is big enough for rates, three secondary scores are worth the effort, each recorded per response as a raw value.

Position. Where the brand appears: first named, in the top three, or later. A brand named in 70% of answers and recommended first in 10% is in a different position from one named in 45% and first in 30%. Report first-named rate separately rather than folding it into a weighted score.

Citation. Whether a page on the brand’s domain is among the cited sources, and which page. In this sample the engines cited vendor documentation, comparison posts and Reddit threads, and rarely the brand’s own homepage; which third-party pages carry a brand is the actionable output, because those are the pages to influence.

Framing. A short categorical code for how the brand is positioned, such as “enterprise”, “budget” or “for existing users of X”. It is more useful than a sentiment score, because the engines are almost never negative and almost always pigeonholing.

Reporting cadence, and what the number cannot tell you

Report monthly, per engine, with the interval, the run-to-run agreement and the sample size on the same page as the rate. If a change is inside the noise floor you measured, say so; the honest sentence is “no detectable change”, not a two-point decline. Refresh the prompt set quarterly, and when you do, keep the old set running for one overlap period so the series is not broken.

AI share of voice is a visibility metric. It says nothing about whether being named produced a visit or a conversion, which is a separate measurement with its own losses, covered in attributing conversions from AI assistants. And it is downstream of a plumbing question: an engine cannot name a page it did not fetch. OpenAI documents that ChatGPT may visit a web page with the ChatGPT-User agent when a user asks a question, and Perplexity documents the same for Perplexity-User, so if your rates are low, the first check is whether those user agents are reaching your content at all, which is what AI crawler log analysis reads from your server. Measure the share of voice, keep the raw responses, and let the interval decide what you are allowed to claim.

Frequently asked questions

How many prompts do I need to measure AI share of voice?

Enough responses per engine for the interval you need. About 100 responses per engine gives a 95% interval roughly 19 points wide on a 50% mention rate, and about 400 brings it to 10 points. Grow the prompt set before adding repeats, because repeats of the same prompt are clustered, and keep at least two runs per period to estimate the noise floor.

Why do ChatGPT and Gemini give different brands for the same prompt an hour later?

Because generative answers are non-deterministic and the engines re-run retrieval each time. In this sample the tracked brands named for the same prompt and engine matched fully in only seven of 16 run pairs, and one ChatGPT prompt went from four tracked brands to none within an hour. That variance is why single-run share of voice numbers are unreliable.

Should AI share of voice be reported as one number across engines?

No. Each engine is a separate population with its own behaviour. In this sample HubSpot appeared in one of 16 ChatGPT responses and five of 16 Gemini responses. Report per engine, and pool only if you also report the per-engine rates.

What is the difference between mention rate and share of voice?

Mention rate is the share of responses that name your brand. Share of voice is your mentions divided by all tracked brands’ mentions, each counted at most once per response. Mention rate is robust to what competitors do; share of voice changes when competitors change, so report both.

How do I know whether a monthly change in AI share of voice is real?

Compare it with the run-to-run variance you measured within the month and with the sampling interval for your sample size. With 100 responses per engine, a change smaller than about 20 points is not distinguishable from noise; with 400, about 10 points. Anything inside that band is “no detectable change”.