InsightsAugust 24, 2026

Prompts for AI visibility tracking: how many phrasings you actually need

How many phrasings does one tracked prompt need when you measure AI visibility? We asked the same question ten different ways across ChatGPT, Gemini, AI Overview and AI Mode. On the chat engines wording is irrelevant, because the model reshuffles anyway. On Google's surfaces it decides everything. 49 runs, with a control group.

By Johannes Gensheimer · 14 min read

Setting up a prompt list for AI visibility tracking means making three decisions: which search intents you cover, how many phrasings you create per intent, and how often you run them. The first comes from your market. The other two are almost always guessed. We measured them.

The usual advice is to use as many phrasings of the same question as you can, because customers word things differently and you would otherwise have blind spots. That sounds sensible, and for one of the four surfaces it is even correct. For two others it is the wrong priority, in a way that systematically skews your tracking.

Because measuring AI visibility means measuring it through prompts, and every prompt is a wording decision. To settle it, we phrased one buying intent ten different ways and sent it through all four surfaces we track: ChatGPT, Gemini, Google's AI Overview and Google's AI Mode. The topic was shopping for project management software, US market, English. 49 runs in August 2026.

Why a naive test does not answer the question

The obvious setup is to run ten phrasings once each and measure the overlap. That produces a number, and the number is uninterpretable.

Because if two rewordings share only 18 percent of their sources, there are two possible explanations: either the rewording changed the answer, or the model reshuffles on every run regardless. With no way to separate those two causes, you have measured something without knowing what.

So one of the ten variants ran three times completely unchanged, not a word different. Whatever differs between those runs cannot possibly be down to phrasing. That is the engine's own noise, and it is the yardstick everything else has to be read against.

Every number here is an overlap between two runs. They differ only in what changed between the runs: in one case nothing, in the other the wording. Only the difference between the two is a finding.

The setup: ten phrasings in three tiers

Ten variants in three tiers, so the result says not just whether rewording costs you something but which kind of rewording does:

  • Word choice, same syntax, different words: best project management software, best project management tool, top project management software, best software for managing projects
  • Sentence form, same words, different grammar: what is the best project management software?, which project management software is best?, recommend the best project management software, what project management software should I use?
  • Conversational padding: I'm looking for the best project management software, what do you recommend? and what's the best project management software to use these days?

One rule mattered throughout: no variant names a vendor. The moment "Asana" appears in the prompt, the vendor analysis is contaminated and the overlap is artificially high.

Comparison across four engines: overlap of cited sources and of recommended vendors, for the same prompt repeated and for reworded prompts
Two rows per engine: sources cited on top, vendors recommended below. The purple dot is the repeat run with an identical prompt, the teal dot the ten rewordings. The distance between them is the actual phrasing effect.

On ChatGPT and Gemini you are mostly measuring randomness

ChatGPT returns only 12 percent shared sources when given the exact same prompt twice. Ten different phrasings return 18 percent. Rewording therefore costs nothing that the next run would not have cost anyway. Gemini shows the same shape at a higher level: 45 percent on repetition, 47 percent on rewording.

Without a control group that would have stayed invisible. "ChatGPT overlaps by only 18 percent when reworded" sounds like a dramatic phrasing effect. It is not one. It is a model that searches fresh and reselects on every run.

For your tracking that means: a single run on these engines is not a measurement, it is a sample of one. Anyone inferring from one ChatGPT result which sources feed "the answer" is reading noise.

On Google, wording is the only variable

The reverse holds on the two Google surfaces. The identical query minutes later returns 100 percent the same sources. Every point lost below that is therefore down to wording alone: 41 points on AI Overview, 59 points on AI Mode.

That is consistent with both surfaces sitting on top of a classic result page. Change the query and you hit a different result page. How a single user question turns into several searches in the first place is something we took apart in our piece on fanout queries.

For most tracking goals the interesting question is not the source mix anyway. It is whether who gets recommended changes. There the picture is much calmer.

Across all four engines and all ten phrasings, Monday, ClickUp and Asana appear in every single answer, Jira in at least nine out of ten. On AI Overview, Monday is the most-mentioned vendor in 8 of 10 phrasings, on AI Mode in 7 of 10.

Even where phrasing demonstrably bites, it bites vendors less than sources: on AI Overview rewording costs 41 points of source overlap but only 18 points of vendor overlap. On AI Mode, 59 against 33.

One exception is instructive. Smartsheet appears in 10 of 10 answers on AI Overview and AI Mode and in 9 on Gemini, but in only 2 on ChatGPT. So the stable core is not the same core everywhere. Within an engine it holds across all phrasings; between engines it does not.

How many of the ten phrasings named each vendor:

Number of the ten phrasings whose answer named that vendor. 10 means: every single one. Shown are all vendors that appeared in at least three of the four engines; the remaining mentions were each confined to a single engine.
VendorChatGPTGeminiAI OverviewAI Mode
Monday10101010
ClickUp10101010
Asana10101010
Jira910109
Smartsheet291010
Trello71097
Notion995

The columns say more than the rows. On AI Overview the list is short and closed, while on ChatGPT one vendor the other three carry throughout, Smartsheet, is almost entirely missing.

Core and long tail: which sources and vendors keep coming back

The same stability can be read from the other side: how many distinct domains and vendors show up across the ten phrasings at all, and how many of those are regulars rather than one-offs?

Distinct domains and vendors across all ten phrasings. The in ≥8/10 column counts the regulars, the once only column the one-offs.
EngineDomains totalin ≥8/10once onlyVendors totalin ≥8/10once only
ChatGPT16082158
Gemini265142595
AI Overview1456960
AI Mode436241252

Two rows stand out. On ChatGPT there is not a single domain that appears in at least eight of ten answers. There simply is no source core to track. And AI Mode draws by far the widest source base at 43 domains, 24 of which appear exactly once. Looking for stable sources there means looking in a very long tail.

On vendors the picture inverts: AI Overview names only 9 distinct vendors in total, and none of them appears just once. You will not find a more closed recommendation list than that.

The four engines recommend the same vendors but cite different sources

The comparison between engines comes out far more decisively than any phrasing effect. Of 65 source domains across the whole experiment, exactly one appeared in all four engines: reddit.com. 21 domains were cited by AI Mode alone. ChatGPT and Gemini overlapped by 3 percent on an identical question.

On vendors the same calculation lands at 38 to 68 percent. The four systems largely recommend the same companies and justify it with almost entirely different sources.

Two matrices across four engines: overlap of cited sources and of recommended vendors on an identical phrasing
Sources on the left, vendors on the right, same colour scale. The right matrix is darker throughout, because the engines agree far more on vendors. The outlined diagonal is not filler: it is the engine against itself, the repeat run, and therefore its own noise.

The diagonal is what makes the comparison readable: ChatGPT agrees with itself 12 percent of the time, while AI Overview and AI Mode agree with each other 42 percent of the time. Two different Google surfaces are more aligned on sources than ChatGPT is with itself.

The practical consequence: a source strategy that works for ChatGPT does not automatically work for AI Mode. And Search Console figures only cover the Google slice of that, as we described in our post on AI search visibility across Google and Bing.

Which rewording actually costs you

Since the phrasing effect only exists on the Google surfaces, the breakdown is only worth doing there. On AI Overview, best project management software, top project management software, best software for managing projects and the two plain questions form a tight block at roughly 74 percent overlap with the rest. Swapping best for top costs nothing.

Two variants fall well out of that block:

  • best project management tool at 43 percent
  • recommend the best project management software at 33 percent

Changing the head noun from software to tool hits a noticeably different result page on Google, even though the intent is identical. So does the imperative. Synonyms in the adjective are cheap; synonyms in the noun are expensive.

Broken down by tier, each cell being the source overlap between variants of the tiers involved:

Mean overlap of cited sources between variants of the tiers involved, across all six possible tier pairings. The Pairs column states how many variant comparisons each mean rests on. Google surfaces only, because that is where a phrasing effect exists at all.
Tiers comparedPairsAI OverviewAI Mode
Word choice ↔ word choice678%43%
Sentence form ↔ sentence form655%43%
Conversational ↔ conversational163%50%
Word choice ↔ sentence form1667%47%
Word choice ↔ conversational848%32%
Sentence form ↔ conversational843%35%

The expensive moves are the two jumps into conversational phrasing. They cost 48 and 43 percent on AI Overview and 32 and 35 percent on AI Mode, in every case well below the comparisons within a single tier. So tracking only keyword-style prompts does not cover how conversational users behave.

One row deserves suspicion: conversational ↔ conversational rests on a single pair of variants, because only two of the ten phrasings fall into that tier. The 63 and 50 percent are therefore one comparison rather than an average, unlike the 6 to 16 pairs behind the other rows.

How to set up your prompt list for visibility tracking

Four rules follow from the numbers, and the third one contradicts the common advice:

  1. One phrasing per intent, not ten. Spend your budget on covering different search intents rather than variants of the same one. The stable vendor core appears in every variant, so extra phrasings buy almost no new information.
  2. Judge visibility at the vendor level, and treat sources separately and more cautiously. "Are we mentioned" is a robust signal. "Through which source" is a shaky one that needs more data points.
  3. On ChatGPT and Gemini you need repetition, not variation. Running the same prompt several times and aggregating beats ten phrasings once each. Anything else measures the model's own noise.
  4. On AI Overview and AI Mode the opposite is true. One run per phrasing is enough, and real variation pays off. Those variants should then match how your audience actually words it, above all in the noun.

In practice: your prompt list is not the same list for all four engines, and your run schedule certainly is not.

The ten variants in detail

For each variant, the mean source overlap with the other nine. The two footer rows supply the yardstick: where an engine's mean sits at its own noise level, every swing in that column is measurement noise rather than a property of the phrasing.

Percentages: mean overlap between this variant's cited sources and those of the other nine, per engine. High means the rewording changes little. Tier A = word choice, B = sentence form, C = conversational padding.
#PhrasingTierChatGPTGeminiAI OverviewAI Mode
1best project management softwareA23%40%74%56%
2best project management toolA28%54%43%25%
3top project management softwareA3%48%74%32%
4best software for managing projectsA16%50%74%56%
5what is the best project management software?B29%50%74%56%
6which project management software is best?B29%46%74%56%
7recommend the best project management softwareB15%38%33%27%
8what project management software should I use?B14%57%49%33%
9I'm looking for the best … what do you recommend?C4%42%49%41%
10what's the best … to use these days?C17%44%45%29%
mean for this engine18%47%59%41%
of which own noise12%45%100%100%

Bold marks the variants that scatter clearly more than is typical for that engine. They occur only in the two Google columns. Nothing is marked in the ChatGPT and Gemini columns, because there the mean sits essentially on the noise level and outliers cannot be told apart from randomness at all.

What this experiment does not show

Publishing numbers without their limits is not something we are willing to do, so here they are:

  • One topic, one market, one point in time. We measured a broad, listicle-saturated tech category in English. Whether a narrow B2B niche behaves the same way is an open question.
  • The 100 percent on Google is short-term determinism. The control runs were about eight minutes apart. How far AI Overview and AI Mode drift across days is not measured here, and real tracking runs daily.
  • ChatGPT cites little. An average of 3.5 domains per answer, against 12.8 on AI Mode. Overlap values jump around on sets that small. The comparison against its own noise floor stays valid, because both sides are computed on equally small sets, but the absolute percentage should not be over-read.
  • The control group is small. Three repeats per engine. That is enough to tell 12 percent from 100 percent, not enough for a precise point estimate.

The obvious next step is the same setup on a narrow non-English niche, with this run as the contrast case.

Until then the short version of the four rules holds: breadth across intents, repetition on the chat engines, and settle whether you get mentioned before asking where it came from. That is exactly what we do with B2B companies, as a GEO agency in Munich, Berlin and Zurich.

Frequently asked questions

For the question of whether your brand gets recommended, one is enough. Across ten phrasings of the same intent the vendors named overlap by 51 to 82 percent, and the stable core appears in every variant. Extra phrasings buy almost no new information there. For source tracking on the Google surfaces several variants do pay off, because they hit genuinely different result pages.

It depends on the engine. On ChatGPT two runs of an identical prompt shared only 12 percent of their cited domains in our experiment, on Gemini 45 percent. There a single run is not a measurement but a sample of one, and you need repetition plus aggregation. On AI Overview and AI Mode repetition adds no new information at all in the short term.

Because the web search runs fresh each time and the model reselects from the hits. Phrasing is innocent here: in our experiment ten different phrasings overlapped at 18 percent, marginally more than two identical runs at 12 percent.

In the short term, yes. In our control group the identical query returned 100 percent the same sources minutes later. That holds for the moment, not across days. How far both drift over longer periods is not something this experiment measures.

On sources, dramatically. Of 65 domains in total, exactly one appeared in all four engines, and ChatGPT and Gemini overlapped by 3 percent on an identical question. On the vendors recommended, agreement runs at 38 to 68 percent. The engines largely agree on who to recommend and almost entirely disagree on where they read it.

Turn visibility into your biggest lead channel

Tell us in two sentences where you stand. You'll get an honest assessment of how much qualified pipeline your visibility could generate within one business day. Free, no strings attached.

Prefer email? Reach us directly at johannes@fento.ai.