· 5 min read
How many prompts does an AI visibility prompt set need?
At least 30, and usually 100 or more. How to size a prompt set from the change you want to detect, the topics you report and the markets you track.
A prompt set needs at least 30 prompts to say anything, and 100 or more if you report a monthly figure and want to know whether it moved. Beyond that, the right size depends on two things: the smallest change you care about, and how many slices, such as topics and markets, you plan to report.
Most teams size their set the other way round, from the number of prompts their tracker's plan includes. That is how sets end up too small to support the conclusions drawn from them.
The short answer
| What you want to do | Prompts you need |
|---|---|
| A rough read of one market | 30 or more |
| A monthly headline figure you can quote | 100 or more |
| Spot a 15-point change at a 25% visibility rate | about 130 |
| Spot a 10-point change at a 25% visibility rate | about 300 |
| Spot a 10-point change at a 10% visibility rate | about 140 |
| Spot a 5-point change at a 25% visibility rate | about 1,200 |
| Report each topic | 5 per topic at the very least, 20 or more to compare them |
| Report several markets | All of the above, per market |
The change figures assume 95% confidence and 80% statistical power, and count each distinct prompt once. The working is in how big a change in AI visibility has to be to be real.
Why the size depends on the change, not the level
Reading a single month's figure is forgiving. Knowing whether it moved is not. At a 25% mention rate, 100 prompts give a margin of about ±8.4 points on one month, but a change between two months has to be around 17 points before you can trust it, because both months carry their own uncertainty.
So start with the question the numbers must answer. If leadership wants to know whether a quarter's work moved visibility by 10 points, a 100-prompt set cannot tell them, whatever the dashboard shows. If they only need to know whether you appear at all in a market, 30 well-chosen prompts will do.
What the research says about the floor
Search engine research faced the same question decades ago: how many test queries does it take to compare two systems reliably? Buckley and Voorhees, working with the TREC evaluations, confirmed the rule of thumb that a good experiment needs at least 25 queries, and 50 is better. A later paper by the same authors showed that even with 50 queries, statistically significant differences between systems can turn out wrong unless the gap is large.
AI visibility tracking is the same kind of measurement: a sample of questions standing in for all the questions buyers ask. The floor of 25 to 50 carries over. It is where a set becomes usable, not where it becomes precise.
Tool vendors have landed higher. Profound's prompt-design guide says its users generally track 100 to 1,000 prompts, that a couple of hundred is typical, and recommends starting with 100.
Why topics and markets multiply the number
Every figure you report is its own sample. A 100-prompt set reported as eight topics is eight samples of about 12 prompts, and at a 25% rate a 12-prompt slice carries a margin of more than ±20 points. The same happens with markets, personas and assistants.
That gives a simple sizing rule: decide the smallest slice you will report, give it enough prompts to stand alone, then multiply.
- Topics. RateMyPrompts' coverage check treats a topic with fewer than five prompts as unreadable. Twenty or more makes large differences between topics visible.
- Markets. Buyers in the UK and the US ask with different constraints, currencies and competitors. Track each market as its own set, rather than one set with a few local prompts mixed in.
- Journey stages. You do not need every stage at full depth, but a stage with two prompts is a guess.
Size is not quality
A bigger set only helps if the extra prompts are different questions a buyer would ask. Three things look like size but are not:
- Paraphrases. "best CRM for agencies" and "top CRM for an agency" are one question. The scorecard counts near-duplicates and template openings for this reason.
- Repeats. Running the same prompt daily adds answers, not information. A month of daily runs on 50 prompts is worth about 70 independent observations. See why more prompts beat more repeats.
- Branded prompts. Three hundred prompts that name your brand measure your reputation precisely and your discovery not at all. See why branded prompts inflate your AI visibility score.
A precise measurement of the wrong questions is still the wrong measurement. Size the set after you have made sure the prompts are written from buyer language.
How RateMyPrompts treats set size
The scorecard grades sets of 3 to 2,000 prompts. It does not reward size directly, because a large set of poor prompts should not outscore a small set of good ones. Instead:
| Set size | What the scorecard says |
|---|---|
| Under 20 prompts | Flagged as too small to detect a change, and confidence shown as low |
| 20 to 49 prompts | Flagged as small enough that only large moves will be readable |
| 50 prompts or more | No size flag; topic depth is checked in coverage instead |
What to do
- Write down the smallest change you need to detect, and your current visibility rate.
- Read the prompt count off the table above.
- Divide by the slices you report. Check each topic has at least five prompts, and ideally 20 or more.
- Give each market its own set of that size.
- Spend on distinct prompts before repeats or extra runs.
- Check the result. The free RateMyPrompts scorecard flags small sets, shows topic depth, and finds the near-duplicates that make a set look bigger than it is.
Questions people ask
- Is 50 prompts enough for AI visibility tracking?
- For a rough read of one market, yes. For reporting change, rarely. At a 25% mention rate, 50 prompts carry a margin of about ±12 points and can only reliably detect a change of about 24 points between two months. That is enough to spot a collapse, not to measure the effect of a campaign.
- How many prompts should I have per topic?
- Five at the very least, because below that a topic's figure is dominated by one or two prompts. Twenty or more per topic is needed before large differences between topics are readable. If you want to report eight topics properly, that alone points to a set of 160 or more.
- Should I use the same prompts in every country?
- Translate the intent, not the words, and track each market as its own set. Buyers in different countries ask with different constraints, currencies and competitors, and mixing markets in one set blurs every figure. Each market needs enough prompts to stand on its own.
- Does running prompts more often mean I need fewer of them?
- No. Repeated answers to the same prompt are highly correlated, so daily runs add far less information than more distinct prompts. A month of daily runs on 50 prompts is worth roughly 70 independent observations. Size the set by distinct prompts.
Sources
- Buckley, C. and Voorhees, E. M. (2000). Evaluating Evaluation Measure Stability. SIGIR 2000
- Voorhees, E. M. and Buckley, C. (2002). The Effect of Topic Set Size on Retrieval Experiment Error. SIGIR 2002
- Lafferty, N. (2026). How to Design Prompts for AI Visibility Tracking in 7 Practical Steps. Profound
- Binomial proportion confidence interval (Wilson score interval). Wikipedia
Grant Simmonds
Director, theround ltd, the company behind RateMyPrompts and the Zebora AI visibility consultancy. More from Grant