GUIDE

Building a Commercial Prompt Universe for a B2B SaaS Product

The buyer questions worth measuring — and how to keep the set honest.

Building a Commercial Prompt Universe for a B2B SaaS Product

The buyer questions worth measuring in AI search are the ones tied to a real buying decision — category, use-case, comparison, integration, security and pricing intent — phrased the way a buyer actually types them into an assistant. Build them into a prompt universe: an open discovery panel you explore freely, plus a small frozen measurement panel (roughly 15–40 questions) you observe repeatedly so results stay comparable. Tag branded, non-branded and competitor prompts separately, and never treat prompt counts as search volume.

Who this is for

This guide is for SEO leads and product marketers who have decided to measure AI search visibility and now need to define what to measure — before observing a single surface.

Use it when you are standing up measurement for the first time, or when your current list is really a keyword export wearing a question mark.

Definitions worth getting straight

  • Prompt universe — the curated set of buyer questions you measure, organized by intent.
  • Discovery panel — an open, evolving list used to find which questions matter. It is allowed to change.
  • Measurement panel — a frozen subset you observe repeatedly. Freezing it is what makes before/after comparable.
  • Branded vs non-branded — branded prompts name you or your product; non-branded prompts describe the category or job. They answer different questions and must be read apart.
  • Controls and treated prompts — controls are questions you do not expect a change to affect; treated prompts are the ones a specific fix should influence. Keeping both lets you tell a real effect from a sitewide shift.

The method: build by intent, then freeze

1. Enumerate the six commercial intents

For each intent, write the questions a buyer asks an assistant during research:

Intent Example question shape
Category “best [category] for [segment]”
Use case “how do teams [job to be done] with software”
Comparison “[you] vs [competitor]”, “alternatives to [competitor]”
Integration “does [category tool] integrate with [system]”
Security “is [category tool] SOC 2 / GDPR compliant”
Pricing “how much does [category] software cost for [segment]”

Decision point: if an intent does not map to a moment in your buyer’s journey, drop it — do not pad the set for symmetry.

2. Phrase like a buyer, not like a keyword tool

Write “What’s the best contract analytics tool for a 20-person legal team?” — not “contract analytics tool”. Assistants respond to natural questions, and buyer phrasing is what surfaces the answers you actually need to measure.

3. Split discovery from measurement

Keep a broad discovery panel you can add to any time. Then choose a frozen measurement panel — the subset you commit to observing on a schedule. Change the frozen set deliberately and rarely, and record when you do, because every change breaks comparability with prior observations.

4. Tag every prompt

Tag each question by intent, by branded/non-branded, by market and language, and mark controls vs treated prompts. Tags are what let you read “we improved on comparison prompts but not category prompts” instead of one blurry average.

5. Size it to what you can actually observe

A measurement panel you can run repeatedly across surfaces beats a giant list you observe once. Sizing is a function of how many surfaces and repeats you can sustain — not ambition.

An honest worked example

A compact starter panel for a contract-analytics product might freeze around these, tagged by intent:

  • Category: “best contract analytics software for mid-market legal teams” · non-branded
  • Use case: “how do legal teams review contracts faster with AI” · non-branded
  • Comparison: “[Product] vs [Competitor A]” · branded
  • Integration: “does contract analytics software integrate with Salesforce” · non-branded
  • Security: “is [Product] SOC 2 Type II compliant” · branded
  • Pricing: “how much does contract analytics software cost per seat” · non-branded

This is an illustrative starter set, not a recommended final panel. Your intents, segments and competitors replace these.

Common failures

  • Keyword dumps in disguise. A list of nouns with “best” bolted on is not a buyer question.
  • All category, no comparison. Comparison and integration prompts are where shortlisting happens; skipping them hides your hardest gaps.
  • Never freezing. If the set changes every week, you can never attribute a change to anything.
  • Untagged prompts. Without intent and branded tags, results collapse into one uninterpretable number.
  • Treating counts as demand. A bigger prompt universe is not more search volume.

Limitations of a prompt universe

A prompt universe describes what is worth measuring, not how many people ask it. It does not estimate demand, and it does not guarantee an assistant will ever be asked your exact question. It is an instrument for comparable observation — nothing more, and nothing less.

Checklist

  • All six commercial intents considered; irrelevant ones dropped on purpose.
  • Questions phrased as a buyer would type them.
  • Discovery panel and frozen measurement panel kept separate.
  • Every prompt tagged by intent, branded/non-branded, market and language.
  • Controls and treated prompts marked.
  • Panel sized to what you can observe repeatedly across surfaces.
  • Prompt counts never presented as search volume.

Where CitePatch fits

Your prompt universe is the input, not the output. The AI Search Radar takes the frozen panel and runs it on a schedule across surfaces, keeping the questions stable so a change in the answer is a signal rather than noise.

If you want to see the shape of the output before building your own set, open the sample report.

Frequently asked questions

What is a prompt universe?
A prompt universe is the curated set of buyer questions a team measures in AI search. It is organized by commercial intent — category, use case, comparison, integration, security and pricing — and phrased the way buyers actually ask, not the way keyword tools phrase things.
How many prompts should we measure?
Start with a frozen measurement panel of roughly 15–40 questions that cover your real buying intents, plus a separate, larger discovery panel you can explore freely. Quality and coverage of intent matter more than raw count; a huge list you cannot observe repeatedly is worse than a focused one you can.
What is the difference between a discovery panel and a measurement panel?
The discovery panel is an open, changing list you use to find which questions matter. The measurement panel is a frozen subset you observe repeatedly so week-to-week results are comparable. Freezing the measurement set is what lets you separate a real change from normal variance.
Are generated prompts the same as search volume?
No. A prompt universe describes the questions worth measuring; it is not a claim about how many people ask them. Treat prompt counts as coverage of intent, never as demand or search-volume figures.
Should we include branded and competitor prompts?
Yes, but label them. Branded prompts (your name, your product) and non-branded category prompts answer different questions, and comparison prompts naming competitors are where buyers make shortlisting decisions. Keep them tagged so you can read each intent separately.
View sample report

Evidence & sources