1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchmeasurement

How Prompt Volumes and Difficulty Scores Are Calculated for AI Search

How we read prompt volume bands and difficulty scores when prioritizing GEO work for clients, and the misreadings we correct most often.

A client asked us last quarter why their "highest volume prompt" was not the top priority in our GEO roadmap. Fair question. The answer required explaining what those volume bars measure, and the explanation improved the whole engagement, so here it is in writing.

We run client programs on Promptwatch, which publishes its methodology, a rarity in this category and the reason we can write this post without guessing. The facts below come from its documentation, checked this week.

The number is a band, on purpose

Prompt volume estimates how often a question is asked of AI models. Nobody outside the model providers measures that directly, so the estimate is built from proxies: keywords attached to the prompt, each carrying monthly volume for Google, Bing, and AI search, refreshed at most once a quarter. The prompt's score is a weighted average across those keywords, with AI search weighted five times, Bing 1.5, and Google 0.5. What you see is a band from 1 to 10k+, drawn as one to five bars, and "No data" until at least one keyword has volume.

The weighting is the part worth pausing on. AI search volume counts five times as much as Google volume in the average, which reflects what the number is for: estimating how often a prompt is put to an AI model, not how often it is typed into Google. Bing sits in the middle at 1.5. Google still contributes, but at half a point of weight per unit of volume, because Google volume is a proxy for demand, not a direct measure of AI asking. The band is capped at 10k+ and drawn as one to five bars, and a prompt shows "No data" until at least one of its attached keywords has any volume at all. That floor matters when you are scanning a long prompt list, because "No data" is not the same as zero demand. It means the proxies have not surfaced a signal yet, which is a different kind of decision than a prompt that genuinely has no volume.

The first misreading we correct: treating the bars as a live trend. The underlying data refreshes quarterly at fastest, so a band is a prioritization signal, not a weekly KPI. When a client wants movement to watch, we point them at visibility and citations, which are measured daily. Volume tells you where to dig; it does not tell you whether digging is working. The two get conflated on a slide because both look like numbers on a dashboard, and the conflation is what sends teams to rewrite pages that were never the bottleneck.

The second misreading: assuming high volume means high priority. This is where difficulty earns its place. Difficulty blends keyword competition in the topic with the authority of the sources AI models already cite for that prompt, over the past 30 days. That second ingredient changes decisions. A five-bar prompt where the models cite government sites and two dominant industry publications is a year-long project. A three-bar prompt cited from a forum thread and someone's outdated blog post might flip in six weeks. We have had more client wins from the second kind than the first. The five-bar prompt feels like the bigger prize on the volume chart, and it is, but the prize is measured in calendar time to move it, and difficulty is the calendar-time estimate.

The reason the second ingredient matters is that AI visibility is not a function of how many people ask. It is a function of who the models already trust on the question. If the models are pulling answers from a government domain and two entrenched industry publications, you are not out-writing them in a sprint. You are building the kind of source the models learn to cite over months, which means link earning, structured coverage, and patience. If the models are citing a forum thread and a stale blog post, the bar is low enough that a well-structured, well-sourced page can displace them in a reasonable window. Difficulty is the score that tells you which situation you are in before you commit the work.

The ranking that does not transfer

The methodology also names the trap we spend the most time talking clients out of: Google position does not predict AI visibility. Models do not rank pages. They compose answers from whichever sources they judge relevant for the specific prompt, so a page that owns position one in Google can go completely uncited, and a page on Google's fourth page can be a model's favorite source. We treat the gap between the two as the opportunity list: prompts where the client already earned Google demand but has zero AI presence.

That gap is the most useful artifact of running both systems together. A page that ranks and is uncited means the content is good enough for a crawler and not yet good enough for a composer, which is a fixable problem. A page that is cited and does not rank means the model found something it liked that the classic SEO stack underweights, which is a signal worth reading. Either way, the gap is where the work is, and the gap only shows up when you stop assuming the two rankings move together.

The two cases call for different work. The page that ranks and is uncited usually needs restructuring for the way models compose: clearer answers near the top, sources and entities named, and the kind of clean factual blocks that a composer can lift. The page that is cited and does not rank is a different lesson. It tells you the model found something the classic SEO stack does not reward, which is often a specific phrasing, a comparison, or a first-person data point. Reading that signal tells you what to make more of, even when Google does not rank the original page. Both cases start from the same observation, which is that the two rankings diverged, and both cases only become visible when you stop treating Google position as a proxy for AI presence.

That is also why we seed prompt lists from the client's own Search Console queries rather than a brainstorm, then let the volume bands and difficulty scores decide the order of attack. Real demand in, published scoring on top, human judgment for the final cut. The workflow is in our post on the GSC import into Prompt Explorer. Google's generative AI performance reports stay useful for Overviews. They do not replace volume bands on assistant prompts.

For budget planning: prompt tracking with volumes, difficulty, and fan-outs starts on the $95/mo Essential plan; agency plans, which is where we operate, start at $199/mo with unlimited prompts.

FAQ

Is the volume band a weekly KPI?

No. The underlying data refreshes quarterly at fastest. A band is a prioritization signal. When a client wants movement to watch, we point them at visibility and citations, which are measured daily.

Does high volume mean high priority?

Not by itself. Difficulty blends keyword competition with the authority of the sources models already cite. A five-bar prompt cited from government sites can be a year-long project. A three-bar prompt cited from a thin blog post might flip in six weeks.

Does Google position predict AI visibility?

No. Models do not rank pages. They compose answers from whichever sources they judge relevant for the specific prompt. We treat the gap between Google demand and zero AI presence as the opportunity list.

If you want a prioritized prompt roadmap instead of a wall of bars, write to hello@1001seomedia.com. We will tell you which of your prompts are six-week wins and which are next year's fight.