Methodology · prompt_fit_rubric@0.2.0

How the Prompt Fit Score works

The Prompt Fit Score grades an AI visibility prompt set out of 100. It weights seven dimensions: buyer relevance (25%), format diversity, branded contamination and topic coverage (15% each), and market fit, long-tail depth and redundancy (10% each). It is deterministic: the same set and brief always get the same score, and no AI model is involved in scoring.

Scale
0 to 100
Dimensions
7, weighted
Set size
3 to 2,000 prompts
AI in the score
None

The question it answers

Does this prompt set measure how your buyers actually ask AI assistants about your category? AI visibility trackers such as Peec, Profound and Semrush run a fixed list of prompts against ChatGPT, Perplexity, Gemini and others, and report how often your brand appears. Every number they report is an average over that list. RateMyPrompts grades the list itself, before anyone trusts the numbers it produces.

It does not measure how visible your brand is. That is the tracker's job. It measures whether the tracker is being asked the right questions.

The seven dimensions

Each scores 0 to 10. Several are measured over the measurable core: unbranded prompts that are about your category rather than a neighbouring one, because variety about the wrong subject is not variety.

DimensionWeightAsksScores highest whenScores lowest when
D1 ICP / commercial relevance25%Are these the questions your buyers ask about your category?45% or more of prompts are unbranded, on-category questions.Capped at 2 when a quarter of prompts use SEO or GEO trade language, and at 3 when 40% or more name your brand.
D2 Format diversity15%Do the prompts ask in the different ways buyers ask?Five or more effective formats among unbranded prompts: rankings, comparisons, conversational questions, personas and so on.A set that is nearly all one shape counts as one format, however many stragglers it has.
D3 Branded contamination15%How much of the set names your own brand?Under 5% of prompts name the brand.40% or more branded scores 2 out of 10.
D4 Taxonomy / node coverage15%Does the set cover several topics, each with enough prompts to read?Five or more topics with at least two prompts each, and none above 40% of the set.One topic holding 70% or more scores 2.
D5 Geographic / language fit10%Do the prompts fit the markets you track?At least 10% of prompts carry a place or currency from your declared markets.20% or more pointing at undeclared markets scores 4.
D6 Stress / long-tail10%Are there specific, constraint-rich prompts?Half or more of on-category prompts run to eight words or more and carry a constraint such as a budget, team size or integration.Under 15% scores 2.
D7 Dedup / redundancy10%How much of the set is paraphrase, or one template repeated?Under 15% of prompts are near-duplicates or share the most common opening.A third or more scores 2.

From dimensions to the headline

The overall score is ten times the weighted sum of the seven dimensions, rounded: round(10 × Σ weight × dimension).

Worked example. Dimensions of 6, 7, 8, 5, 7, 4 and 8 give 10 × (0.25×6 + 0.15×7 + 0.15×8 + 0.15×5 + 0.10×7 + 0.10×4 + 0.10×8) = 10 × 6.4 = 64, Mixed, grade C.

Confidence is shown beside the score and does not change it: low under 20 prompts, medium without a brand and category in the brief, high otherwise.

BandScoreGradeWhat it means
Excellent85 to 100AMeasures buyers well. Keep it stable and track it.
Strong70 to 84BTrustworthy, with a few fixes worth making.
Mixed50 to 69CUsable, but some numbers it produces will mislead.
Weak30 to 49DUnlikely to reflect buyers. Rebuild before reporting.
Critical0 to 29FMeasures something other than discovery.

Beside the score, never in it

Red flags call out problems a person should see whatever the score: personal data in prompts, sets too small to detect a change (under 20 prompts), leading prompts that assume the answer, year-stamped prompts and heavy competitor lock-in. Where a flag's cause is already a dimension, it says so rather than counting twice.

Coverage shows how much of the category's buying decision the set reaches, on a grid of topics and six journey stages: discovery, problem, evaluation, comparison, validation and post-purchase. It is scored out of 100 and banded complete (80+), broad (60+), partial (35+) or thin. It sits beside the headline because a coverage number is partly a score of the grid, and the grid is generated, so it is printed in full and can be edited.

Coverage partPointsMeasures
Topics reached30Share of the category's topics with at least one prompt
Depth on what matters25Share of high-weight topics with five or more prompts
Stages occupied20Share of the six journey stages reached
Balance15Spread across topics and stages, cut when one stage dominates
Staying on the grid10Less the share of prompts that fit no topic

Recommendations rank the changes worth making. Those the tool can make itself, such as trimming branded prompts or dropping near-duplicates, show the points they are worth, worked out by re-running the scorer on the changed set rather than estimated. Re-scoring applies the ones you tick and grades the result as a new, linked scorecard.

What it does not measure

  • How visible your brand is. That is what the tracker reports.
  • How many people ask each question. No auditable volume figure exists for AI prompts, and the score does not invent one.
  • Where your prompts came from. The brief lets you declare it, and it is shown, but it is never scored.
  • Meaning, beyond what rules can detect. Two prompts sharing words can still ask different things; redundancy is measured on wording.

Questions about the method

Does RateMyPrompts use AI to score prompt sets?
No. The Prompt Fit Score is deterministic: the same prompt set and brief always get the same score, and no AI model is involved in scoring. A model may help build the reference grid for coverage, which sits beside the score and is labelled when it does.
Why is buyer relevance weighted highest?
Because it decides what the visibility number means. A set that asks category jargon or names your brand returns more mentions than buyers actually see, so buyer relevance carries 25% and caps apply when jargon or branded prompts dominate.
Why doesn't the score reward prompts that contain the category name?
Real buyers usually describe their situation rather than naming the category. In one live 156-prompt mortgage set, only 13 of 94 unbranded prompts contained the word mortgages. Scoring the keyword would reward keyword-shaped sets, which are the ones that inflate visibility.
How is the rubric checked?
Against prompt sets scored by hand. Tests assert that no set a person rated Critical or Weak scores 50 or more, and no set rated Strong or Excellent scores below 70. Every scorecard records the rubric version it was graded against.

See it on your own set

Free, in under a minute. Terms are defined in the glossary.

Score my prompts

Published · Rubric prompt_fit_rubric@0.2.0