Skip to content
Services
SEOGoogle AdsRemarketingSocial Media AdvertisingContent MarketingConsultingAI VisibilityDigital Advertising
deenes

Resources · AI Visibility

Which prompts should you track?A working method for prompt research

By Philipp Enders · Published August 2026 · 8 min read

Hierarchy of evidence for prompt research: first-party data above keyword tools above synthetic generation

Most AI visibility tools will offer to generate your prompt list for you. Accept the offer and you'll get fifty plausible-looking questions in about four seconds, most of them variations on “best [category] tool for [industry]”. Import them, and the tool will faithfully measure your visibility against a list nobody thought about. Our own AI visibility tracker offers suggestions too — and the same rule applies there: suggestions are orientation, not evidence.

That list is the experiment. Everything downstream — the mention rate, the citation counts, the competitive share chart your client sees — inherits whatever bias went into choosing it. A tracker has no opinion about whether the thing it's measuring is worth measuring. That judgment is the job, and it's the part that gets skipped. Here's how we approach it, including the parts that went wrong on a recent project.

Start with what people actually did, not what a tool estimates

There's a hierarchy of evidence for prompt research, and it's worth being explicit about it because most guides treat all sources as equivalent. First-party behavioural data is the strongest: Search Console queries, site search logs, support tickets, sales call notes. These record what real people typed or said, unmodelled. Seer Interactive have argued this point hardest, from user testing sessions where they watched people actually prompt — their finding is that prompt inventories built from customer behaviour look meaningfully different from ones built from marketer intuition. Third-party keyword tools come second: Semrush, Ahrefs and the rest estimate volume from clickstream panels and modelling — useful for scale and for phrasing, weaker as ground truth. Synthetic generation — asking an LLM what your customers would ask — comes last. It's fine for orientation in an unfamiliar vertical. It's not evidence.

The reason to rank them is that they disagree, and when they do you need to know in advance which one wins. On a recent project for a regional energy supplier, they disagreed badly. The client's Search Console showed strong local demand: thousands of impressions on queries combining a product with the city name. Semrush, running the same seeds with location modifiers, returned nothing. Not low volume — no volume, across every combination I tried. Across 2,255 unique keywords in the exports, exactly three contained a local modifier. That looks like a finding until you check it against the first-party data, where the same demand is sitting there in observed clicks. The German index simply thins out in the local long tail for that vertical. If I'd taken the tool at face value, I'd have dropped the location dimension from a study where 70% of non-brand clicks carried a place name. Observed behaviour beats modelled volume. Write that rule down before you start, because in the moment a confident-looking zero is persuasive.

Your keyword data probably has no questions in it

Here's a trap worth knowing about. The standard advice is to pull W-questions from Search Console. On that same project I ran the check and found zero — not a handful, zero — question-form queries in the top 1,000. The reason is mundane: the GSC export truncates, in that case at eight clicks, and question phrasing lives in the long tail below the cut. There's a second reason too: in commercial categories, question-shaped demand has been migrating to comparison sites and, increasingly, to the chat window itself. It never touches your property, so it can't appear in your data. So first-party data gives you topics, vocabulary, priority and geography. It usually won't hand you the question. For that you go to the keyword tool's question filter, and you accept that you're translating rather than transcribing. One caveat on that translation: when I pulled the question exports, 2,151 of the 2,255 keywords were question-shaped, but only 108 carried commercial or transactional intent. That ratio matters, because the informational questions are the easy ones to fill a tracking list with, and they're the ones least likely to produce a brand recommendation.

Four prompt types, and the one everyone overloads

Sort your candidates by what they're designed to surface:

  • Situational advice: first person, one concrete constraint. “I have X but my situation is Y — what should I do?” Tests whether your content gets retrieved at all.
  • Comparison: X versus Y for a specific case. Tests whether you appear when alternatives are being weighed.
  • Process and prerequisite: “What do I need before I can do X?” Tests your help centre and service content.
  • Recommendation: “Which provider should I choose?” The only type that reliably forces the model to name brands.

Audit a typical tracking account and you'll find most prompts are some version of type 4 phrased as type 1, or four variations of the same category question. Similarweb make the point that comparison and constraint-heavy prompts are where engines actively weigh alternatives and name names — and that these have low or no volume in traditional keyword tools. Both halves of that matter: they're the valuable ones, and you can't find them by sorting a keyword export by volume. Which means you build them by pattern. That's not a weakness in the method as long as you say so.

Two prompt shapes, because two kinds of people are typing

I used to write every test prompt as a full sentence with context. That's wrong, and the survey data is fairly clear about it. Surveys by Stella Rising found around two-thirds of respondents write prompts of fifteen words or fewer, and only about one in eight goes beyond 30 words at all — the elaborate templates that circulate on LinkedIn remain the exception. Semrush's clickstream analysis of ChatGPT's search mode puts average prompt length between four and nine words. At the same time, a growing share of prompts carry personal context — budget, location, life stage, profession — and those are exactly the prompts where models start recommending specific brands. So both are real, and they test different things. Split your set: some themes rendered short and keyword-shaped, close to the original query; others written out with one concrete constraint. Vary the shape between themes, never within a matched pair, or you lose the ability to attribute a difference to anything.

Choosing locations without guessing

If you serve a geography, the location dimension is not optional. On the energy project, 254 of 449 non-brand queries carried a place name, accounting for 70% of non-brand clicks. Local intent was the norm, not a segment. The question is which places. Don't pick from a map — count them in your own data, then build tiers that each answer a different question:

  • No location: your baseline. What does the model say with no geographic signal?
  • Primary city, the one with real volume: does the model recognise you locally?
  • District, or “near me” with no place named: how granular does the resolution go — and does it work when the user assumes rather than states?
  • Surrounding towns: does the model know where your service area actually ends? For a utility this is the interesting one, because the answer is frequently no.
  • National: a control. If a nationwide framing returns comparison portals and large suppliers, that tells you what your local result was worth.

One methodological point that's easy to miss: three separate variables are in play — where the prompt is run, which location is named in the wording, and which language is used. Ethan Lazuk has written the clearest treatment of this. If you're testing from an office in the city you're studying, your “no location” condition isn't location-free: the model can infer position from IP or account. During this project a signed-out Google search reported my location unprompted, based on prior activity. Pick one execution location, hold it constant, and disclose it.

Keep brand names out, with one exception

A prompt containing your brand name will return your brand. Mixing those into a category-level set inflates every number you report. SE Ranking make this point and it's worth being strict about: track branded prompts, if you track them at all, as a separate group with its own baseline. This gets awkward when a client's agreed topic list includes their own product names, which happens more often than you'd think. Rewrite them generically — “charging tariff covering home and public charging on one card” instead of the product name. If someone insists on testing the brand version, it goes in its own bucket and never into the headline figure.

How many prompts, and what your number is actually worth

The unit isn't the prompt. It's the prompt-run: one scored answer, one engine, one prompt, one point in time. Repeated sampling of an identical prompt shifts results by 10–34%. A variance decomposition of nearly 13,000 responses attributed around a third of total variance to resampling the same prompt, and roughly 1.5% to brand identity — the thing you're trying to measure. AirOps' research with Kevin Indig (Growth Memo), across 16,851 queries, found that after running the same prompt three times in ChatGPT, only 2.3% of page citations persisted across all three runs. The practical consequence: if you run fifty prompts once per engine, your mention rate carries a margin of error somewhere around fifteen to twenty points. That's a directional finding, not a metric. Report it as one.

Breadth beats depth: 150 prompts × 3 runs lands near ±5 points, 20 prompts × 30 runs near ±15 — same budget
Same budget, three times the precision: if you have budget for runs, spend it on breadth first.

If you do have budget for runs, spend it on breadth first. MaxAEO's arithmetic on this is worth internalising — 150 prompts run three times lands near ±5 points, while 20 prompts run thirty times costs the same and lands near ±15. Same budget, three times worse.

Two traps on the way out

Your own group counts as you. When I pulled the organic competitor list for that project, the top result by relevance was the client's own network subsidiary, followed by their parent company and their foundation. Anyone building a competitor coding list off that export without checking corporate structure would count the client as their own competitor. Build the coding list before you read a single answer, and mark the entities that belong to the client. And: product-adjacent seeds return nothing. Two of my fourteen seeds came back empty because they were closer to product names than to how people search. That's not absent demand, it's the wrong word. Check against first-party data before you drop a theme.

Then report what your sample size can actually support. A well-chosen fifty prompts run once will tell you honestly whether you show up, which pages get pulled, and who appears instead. It won't tell you your visibility is 32%. Anyone reporting that number from a single pass is reporting noise with a decimal point.

This research layer sits before every tracking setup we run — it's the uncomfortable, indispensable part of our AI visibility work. If you want to know how your brand shows up in AI answers — and what the number behind it actually supports — talk to us.

Prompt research FAQs

For an honest start, around 50 well-derived prompts per engine — read as a directional finding. If you need precision, spend on breadth first: 150 prompts run three times each land near ±5 points of error; 20 prompts run thirty times cost the same and land near ±15.

For orientation, yes; as the basis of measurement, no. Generated lists are mostly generic category questions. Topics and geography belong to first-party data — Search Console, site search, support tickets — and phrasing to keyword tools.

Only as a separate group with its own baseline. A prompt containing your brand name returns your brand — mix it into the category set and every number you report is inflated.

Because models don’t answer deterministically: the same prompt shifts results by 10–34% on repetition. The unit is therefore the prompt-run, not the prompt — and a single pass is a directional finding, not a metric.