· 5 min read
How big does a change in AI visibility have to be to be real?
Most month-on-month AI visibility changes sit inside the margin of error. How to work out the smallest change your prompt set can actually see.
A change in AI visibility is real when it is bigger than the smallest change your prompt set can detect. For most sets that threshold is far larger than the movements people report. A dashboard showing visibility up from 22% to 27% this month is, on a 100-prompt set, showing you noise.
This piece shows how to work out your own threshold, why comparing two months is harder than reading one, and what else moves the number when nothing in your marketing has changed.
Why does a visibility rate have a margin of error?
Because your prompt set is a sample. A tracker asks a fixed list of questions and reports the share of answers that mention you. A different, equally reasonable list of questions about your category would give a different share. The margin of error describes how far the reported figure could sit from the rate you would get across every question buyers actually ask.
The standard way to put a range on a proportion is the Wilson score interval. It behaves well at the small sample sizes and low rates common in AI visibility, where the simpler textbook formula goes wrong. The table below uses it, with n as the number of distinct prompts, each counted once.
How precise is a single month's figure?
| Distinct prompts | Margin at 25% visibility | Margin at 10% visibility |
|---|---|---|
| 20 | ±17.8 pts | ±13.7 pts |
| 30 | ±14.9 pts | ±11.1 pts |
| 50 | ±11.7 pts | ±8.5 pts |
| 100 | ±8.4 pts | ±6.0 pts |
| 200 | ±6.0 pts | ±4.2 pts |
| 500 | ±3.8 pts | ±2.6 pts |
| 1,000 | ±2.7 pts | ±1.9 pts |
Read it as: with 100 prompts and a reported 25%, the true rate across buyer questions is plausibly anywhere from about 17.5% to 34.3%. That is the level. Detecting a change is harder.
Why is a change harder to see than a level?
Because both months have a margin of error, and they combine. To say a change is real you also want a reasonable chance of spotting it when it happens, not just protection against false alarms. The usual standard is 95% confidence with 80% statistical power. The smallest change that meets both, the minimum detectable change, is:
Minimum detectable change = (1.96 + 0.84) × √(2 × p × (1 − p) ÷ n)
where p is the visibility rate and n the number of distinct prompts.
| Distinct prompts | Smallest real change at 25% | Smallest real change at 10% |
|---|---|---|
| 30 | 31.3 pts | 21.7 pts |
| 50 | 24.3 pts | 16.8 pts |
| 100 | 17.2 pts | 11.9 pts |
| 200 | 12.1 pts | 8.4 pts |
| 500 | 7.7 pts | 5.3 pts |
| 1,000 | 5.4 pts | 3.8 pts |
Turned around: at a 25% rate, seeing a 15-point change needs about 130 prompts, a 10-point change about 300, and a 5-point change about 1,200. At a 10% rate the figures are about 60, 140 and 570. Few prompt sets are that large, which is why so many reported gains do not survive a second look.
Don't your daily runs make it more precise?
Much less than the answer count suggests. Most trackers run each prompt once a day, so a month of a 50-prompt set produces 1,500 answers. But answers to the same prompt are strongly correlated, so they are worth far fewer independent observations. At a typical correlation of 0.7 those 1,500 answers are worth about 70, which gives a margin of about ±9.9 points and a smallest real change of about 20 points. The full working is in why more prompts beat more repeats.
Repeats do help with one thing: the order of brands. In a 2025 study of repeated LLM evaluations, a second run removed about 83% of the rank inversions a single run produced. So two runs make "are we ahead of Globex?" steadier, even though they barely narrow the margin on the rate itself.
Why are topic and market figures so much worse?
Because each segment is a smaller sample. A 100-prompt set split across eight topics has about 12 prompts per topic, and at 25% visibility a 10-prompt segment carries a margin of about ±24 points. Topic-level charts on a set that size are mostly noise, however confident the colours look.
The Prompt Fit Score's coverage check treats a topic with fewer than five prompts as unreadable for exactly this reason, and five is a floor, not a target. If you want to compare topics or markets, each one needs enough prompts to stand on its own, and every market is best tracked as its own set.
What else moves the number?
Model updates, often by more than any campaign. When Chen, Zaharia and Zou compared two versions of GPT-4 three months apart, accuracy on one task fell from 84% to 51%, and other tasks moved in the opposite direction. An assistant's model or retrieval changing can shift every prompt at once, and nothing inside your data will tell you it happened.
Treat a known model or version change as a break in the series. Mark it on the chart, start a new baseline, and compare like with like either side of it.
Spreading across assistants helps more than repeating. In a 2026 analysis of 12,933 AI answers about 20 brands, adding models and languages reduced measurement error far more than extra repeats of one prompt did. Report each assistant separately too, because they can move in opposite directions.
What to do
- Work out your smallest real change before the next report. Count distinct prompts, take your current visibility rate, and read the figure off the table above, or use the formula.
- Report the range, not just the number. "25%, plausibly 17% to 34%" is honest; "25%" alone is not.
- Set the threshold in advance. Agree that changes smaller than your minimum detectable change are reported as "no clear change".
- Only chart segments that can stand alone. Five prompts per topic is the bare minimum; 20 or more makes large differences readable.
- Mark model updates and restart the baseline after them.
- If the change you care about is smaller than your threshold, grow the set. More distinct, buyer-shaped prompts is the only thing that narrows it much. See how many prompts a prompt set needs.
The free RateMyPrompts scorecard flags sets too small to detect a change and shows how the prompts spread across topics, so you know what your numbers can support before you report them.
Questions people ask
- Is a 5-point rise in AI visibility significant?
- Only with a very large prompt set. At a 25% mention rate you need about 1,200 independent prompts to detect a 5-point change between two periods with 95% confidence and 80% power, and about 570 at a 10% rate. On a typical 50 to 200 prompt set, a 5-point rise is indistinguishable from noise.
- How do I calculate the margin of error on my AI visibility score?
- Treat the mention rate as a proportion and use the Wilson interval with your number of independent observations, which is roughly your number of distinct prompts. For a quick estimate, the margin is about 1.96 times the square root of p times (1 − p) divided by n. At 100 prompts and 25% that is about ±8.5 points.
- Why does my AI visibility jump around from month to month?
- Three reasons stack up. The set is a sample of questions, so the figure has a margin of error. Answers vary between runs even at temperature zero. And model updates can move every prompt at once. On a small set, those alone produce swings of 10 points or more with nothing changing in your marketing.
- Does tracking more AI assistants make the number more precise?
- Usually more than repeating the same prompts does. A 2026 study of AI answers about brands found that adding models and languages reduced error far more than extra repeats. Report each assistant separately as well, because they can move in opposite directions.
Sources
- Binomial proportion confidence interval (Wilson score interval). Wikipedia
- Power of a test. Wikipedia
- Chen, L., Zaharia, M. and Zou, J. (2023). How is ChatGPT's behavior changing over time? arXiv 2307.09009
- Alvarado Gonzalez, M. A. et al. (2025). Do Repetitions Matter? Strengthening Reliability in LLM Evaluations. arXiv 2509.24086
- Żatuchin, D. (2026). Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers. arXiv 2607.13304
Grant Simmonds
Director, theround ltd, the company behind RateMyPrompts and the Zebora AI visibility consultancy. More from Grant