Tools / Free calculator

AI visibility sample-size calculator

Work out how many prompts, runs and assistants a prompt set needs before a change in your AI visibility is real rather than noise. Enter your set, the assistants you track and the change you care about. The calculator returns the noise band, the smallest change one round can detect, and how long you would wait to see yours.

Plan your measurement

Start with your own numbers. The defaults are a typical 50-prompt set on one assistant, with a 25% visibility rate and a 5-point change worth reacting to.

Assistants tracked (1 selected)

Each assistant answers every prompt, so more assistants means more answers per round.

Run-to-run stability

No cadence rescues a set this size: you would need 24 rounds, about 24 months, to see a 5-point move. Running it more often will not help. Grow the set to roughly 1,178 prompts, or track a change of 25 points or larger.

Noise band, one round
±11.7 pts
Smallest real change
24.3 pts
Effective sample
50
of 50 responses

More prompts beat more repeats

Running the same prompt again mostly re-measures the same thing, so repeats are discounted. Smaller bars are better.

  • As designed50×124.3 pts
  • Triple the runs50×321.7 pts
  • Triple the prompts150×114.0 pts
How this is calculated

Visibility is a proportion, so the interval is a Wilson score interval, which stays inside 0 and 100 at small samples and extreme rates where the usual normal approximation does not.

Repeated runs of one prompt are a cluster, not independent samples. The effective sample is the response count divided by the design effect, 1 + (repeats − 1) × correlation. At 0.70 correlation your 50 responses carry about 50 independent responses of information.

The smallest real change is the minimum detectable effect for two proportions at 95% confidence and 80% power. Pooling rounds multiplies the effective sample, so the detectable change falls with the square root of the number of rounds — hence 24 rounds, about 24 months.

These are design calculations on the assumptions above, not an analysis of your actual runs. They assume prompts are a fair sample of the questions you care about and that the model mix is stable between rounds. A model version change breaks that: treat it as a new baseline rather than a movement. Tracking many segments at once also inflates false positives; a five-point move in one topic out of twenty is weaker evidence than the same move overall.

Why this tool exists

AI visibility trackers report a single number, usually with no margin of error. But the answers behind it vary from run to run, even with the model set to be deterministic: one study found accuracy moving by up to 15% between identical runs, and another got 80 different answers from 1,000 identical requests. Researchers now recommend describing AI visibility as a distribution, not a single reading.

So before anyone reads a rise or fall as the result of a campaign, one question comes first: can this prompt set, run this way, see a change that size at all? For most sets, the honest answer is no. This calculator answers it in advance, so you size the set to the question rather than the tracker plan.

How the calculation works

  1. 1. Effective sample

    Every prompt answered by every assistant, on every run, is one answer. But answers to the same prompt are strongly alike, so they are discounted by the design effect: 1 + (repeats − 1) × correlation, where repeats is runs × assistants. The effective sample is the number of answers divided by it. Treating each assistant like a repeat is cautious; evidence suggests different models add a little more information than that.

  2. 2. Noise band

    The visibility rate is a proportion, so its 95% range comes from the Wilson score interval, which stays accurate at small samples and low rates where the textbook formula breaks down.

  3. 3. Smallest real change

    Comparing two rounds combines two margins. The smallest change you can trust, with 95% confidence and an 80% chance of spotting it when it happens, is (1.96 + 0.84) × √(2 × p × (1 − p) ÷ n), with p the visibility rate and n the effective sample.

  4. 4. How long to wait

    If one round cannot see your change, the calculator works out how many rounds you would need to combine, and how long that takes at your cadence. It assumes the rounds are comparable; a model update between them breaks that.

Why this matters for honest experiments

Every content change, PR push or new page is an experiment on your AI visibility. Experiments go wrong in predictable ways, and this calculator is there to stop the common ones:

  • Decide the threshold before you look. Agree the change worth acting on, and treat anything smaller as no clear change.
  • Don't keep checking. Every extra look at a noisy number is another chance of a false alarm.
  • Beware of slicing. Across 20 topics, one will move 5 points by chance. A segment needs its own sample.
  • Reset after model updates. One study saw a model's accuracy on a task fall from 84% to 51% between versions three months apart. Treat a version change as a new baseline.
  • Buy prompts before repeats. Distinct prompts narrow the noise band far faster than running the same prompts again.

Worked examples

At a 25% visibility rate and typical run-to-run stability (correlation 0.7).

DesignAnswersEffectiveNoise bandSmallest real change
50 prompts, run daily on ChatGPT, read monthly1,50070±9.9 pts20.4 pts
100 prompts, once each on 3 assistants300125±7.5 pts15.3 pts
250 prompts, once each on 1 assistant250250±5.3 pts10.9 pts
500 prompts, twice each on 3 assistants3,000667±3.3 pts6.6 pts

The first row is the common case: 1,500 answers a month that carry the information of about 70. Read more in why more prompts beat more repeats and how big a change has to be to be real.

Questions

How many prompts do I need to detect a 5-point change in AI visibility?
At a 25% visibility rate, about 1,200 prompts each answered once, for a change between two rounds at 95% confidence and 80% power. At a 10% rate it is about 570. Most prompt sets are far smaller, which is why small month-on-month moves are usually noise.
Does tracking more AI assistants make the result more precise?
Yes, and the calculator treats each extra assistant like an extra run of every prompt, discounted for how alike the answers are. That is cautious: a 2026 study of AI answers about brands found that spreading across models reduced error more than repeating one prompt.
What run-to-run stability should I choose?
Fairly stable (a correlation of 0.7) is typical: whether your brand appears is mostly decided by the prompt, not the run. Choose varies a lot if your tracker shows answers changing often between runs, and very stable if they barely change.
Is this an analysis of my tracking data?
No. It is a design calculation on the numbers you enter, done before you collect data. It tells you what a given set and cadence can detect, so you can size the set or set expectations before anyone reads a chart.

Sources

  1. Binomial proportion confidence interval (Wilson score interval). Wikipedia
  2. Kish, L. (1965). Survey Sampling. Summarised in "Design effect", Wikipedia
  3. Power of a test. Wikipedia
  4. Schulte, Bleeker and Kaufmann (2026). Don't Measure Once: Measuring Visibility in AI Search (GEO). arXiv 2604.07585
  5. Żatuchin (2026). Where Does the Noise Come From? arXiv 2607.13304
  6. Atil et al. (2024). Non-Determinism of "Deterministic" LLM Settings. arXiv 2408.04667
  7. He and Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference
  8. Chen, Zaharia and Zou (2023). How is ChatGPT's behavior changing over time? arXiv 2307.09009