Blog5 min read

How many AI answers do you need? Sample size and margin of error for AI visibility

Concepts
A chart of the margin of error of an AI visibility rate, falling from 31 points at 10 answers to 3 points at 1,000 answers

A mention rate from 30 AI answers carries a margin of about 18 points. How many prompts, engines and runs it takes to measure AI visibility and to compare two periods, with the math.

Why AI visibility needs a sample

An AI answer is a draw. The same buying question gets different brands, in a different order, from different sources from one run to the next. A mention rate is the share of answers that name your brand, so like any measured share it carries sampling error: the rate you measured sits near the true rate, and the number of answers sets how near.

The margin of error gives that distance. At 95% confidence, a rate p measured from n answers sits within this many points of the true rate:

margin = 1.96 * sqrt(p * (1 - p) / n) * 100

The margin is widest when the rate is near 50%, so the table below plans for 50%. At a rate of 20% every margin is a fifth smaller.

Margin of error by number of answers

Margin of error at 95% confidence for a rate near 50%
AnswersMargin of errorExample
10±31 pointsOne prompt on one engine, run 10 times
21±21 pointsThe smallest AI Index category, US running shoes
30±18 pointsThe floor monitors flag with quality.lowSample
71±12 pointsOne engine in Korea in the AI Index
100±10 points20 prompts × 1 engine × 5 runs
200±7 points20 prompts × 2 engines × 5 runs
400±5 points20 prompts × 4 engines × 5 runs
479±4.5 pointsEvery AI Index answer, all engines and markets
1,000±3 points20 prompts × 5 engines × 10 runs

The AI Index shows the table at work. Korea's ChatGPT rate of 61% rests on 71 answers, and its interval runs from 49% to 71%. The 82% across all 479 answers is far firmer, at 79-85%. Category rankings rest on 21 to 48 answers each, so two brands a few points apart in one category are tied.

Rates near 0% and 100%

At the edges the simple formula shrinks the margin to zero. Zero mentions in 30 answers gives a margin of zero, yet a brand named in 5% of all answers draws zero mentions in one sample of 30 out of five.

The Wilson interval holds at the edges, and the AI Index publishes it. Zero mentions in 30 answers has a 95% interval of 0-11%. On was named in all 21 US running shoe answers, an interval of 85-100%.

Comparing two periods

A change between two periods carries the noise of both. With 100 answers in each period, two periods with the same true rate differ by up to 14 points in 95 cases out of 100. A 10-point swing between two weeks of 100 answers each sits inside that range.

Answers in each period to detect a change
Change to detectAnswers in each period
25 pointsabout 65
20 pointsabout 100
15 pointsabout 175
10 pointsabout 400
5 pointsabout 1,600

These counts detect the change 80% of the time at 95% confidence when the rates sit near 50%. A monitor compares a window with the one before it on the same prompts and engines, and reports the change when the two periods are comparable. quality.comparable turns false when a period has no answers or when the scoring rules changed between them.

Where the answers come from

Answers in a period = prompts × engines × runs. A monitor with 20 prompts on ChatGPT, Gemini and Perplexity, run daily, collects 60 answers a day: 420 a week, 140 of them per engine. That is ±8 points per engine each week and ±4 points per engine over 30 days.

ChatGPT takes 2 credits a task and Gemini and Perplexity 1 each, so that monitor costs 80 credits a run, about 2,400 credits a month on a daily schedule. Pricing lists the credits each plan includes.

  • Spread the sample over prompts. Answers to one prompt resemble each other, so adding prompts firms up a rate faster than running the same prompts more often.
  • Read the overall rate first, then each engine, then single prompts. Each slice has fewer answers and a wider margin.
  • Set the window long enough to reach your target count: a week for large monitors, 30 days for small ones.
  • Hold prompts, engines and aliases fixed across the periods you compare, so both periods measure the same thing.

How to report AI visibility numbers

  • Give the count with the rate: “61% of 71 answers” tells the reader how firm the number is.
  • Add the interval when a decision rests on the number: “61% (49-71%)”.
  • Read gaps smaller than the margin as ties, between brands and between periods.
  • Open the answers behind a surprising number to see which prompts and engines moved it.

Frequently asked questions

How many answers do I need to measure AI visibility?

At least 30 scored answers per window for each rate you read, which gives a margin of about ±18 points. 100 answers bring the margin to ±10 points and 400 to ±5. Comparing two periods takes about 100 answers per period for a 20-point change and about 400 for a 10-point change.

How many prompts should a monitor track?

Enough that prompts × engines × runs reaches your target count in each window. 20 prompts on three engines, run daily, give 140 answers per engine a week. Adding prompts firms up a rate faster than running the same prompts more often.

Why did my mention rate change this week?

Part of every change is sampling noise. With 100 answers a week, two weeks with the same true rate differ by up to 14 points in 95 cases out of 100. Compare longer windows or add prompts, and open the answers behind the change.

What confidence level does querying.ai use?

95%, the common standard. The AI Index publishes 95% Wilson intervals, and monitors publish the sample size of every window.

Start free

2,000 free credits when you sign up. Collect your first answers in minutes.

Start free

More from the blog