Setting up a prompt list for AI visibility tracking means making three decisions: which search intents you cover, how many phrasings you create per intent, and how often you run them. The first comes from your market. The other two are almost always guessed. We measured them.
The usual advice is to use as many phrasings of the same question as you can, because customers word things differently and you would otherwise have blind spots. That sounds sensible, and for one of the four surfaces it is even correct. For two others it is the wrong priority, in a way that systematically skews your tracking.
Because measuring AI visibility means measuring it through prompts, and every prompt is a wording decision. To settle it, we phrased one buying intent ten different ways and sent it through all four surfaces we track: ChatGPT, Gemini, Google's AI Overview and Google's AI Mode. The topic was shopping for project management software, US market, English. 49 runs in August 2026.
Why a naive test does not answer the question
The obvious setup is to run ten phrasings once each and measure the overlap. That produces a number, and the number is uninterpretable.
Because if two rewordings share only 18 percent of their sources, there are two possible explanations: either the rewording changed the answer, or the model reshuffles on every run regardless. With no way to separate those two causes, you have measured something without knowing what.
So one of the ten variants ran three times completely unchanged, not a word different. Whatever differs between those runs cannot possibly be down to phrasing. That is the engine's own noise, and it is the yardstick everything else has to be read against.
Every number here is an overlap between two runs. They differ only in what changed between the runs: in one case nothing, in the other the wording. Only the difference between the two is a finding.
The setup: ten phrasings in three tiers
Ten variants in three tiers, so the result says not just whether rewording costs you something but which kind of rewording does:
- Word choice, same syntax, different words:
best project management software,best project management tool,top project management software,best software for managing projects - Sentence form, same words, different grammar:
what is the best project management software?,which project management software is best?,recommend the best project management software,what project management software should I use? - Conversational padding:
I'm looking for the best project management software, what do you recommend?andwhat's the best project management software to use these days?
One rule mattered throughout: no variant names a vendor. The moment "Asana" appears in the prompt, the vendor analysis is contaminated and the overlap is artificially high.

On ChatGPT and Gemini you are mostly measuring randomness
ChatGPT returns only 12 percent shared sources when given the exact same prompt twice. Ten different phrasings return 18 percent. Rewording therefore costs nothing that the next run would not have cost anyway. Gemini shows the same shape at a higher level: 45 percent on repetition, 47 percent on rewording.
Without a control group that would have stayed invisible. "ChatGPT overlaps by only 18 percent when reworded" sounds like a dramatic phrasing effect. It is not one. It is a model that searches fresh and reselects on every run.
For your tracking that means: a single run on these engines is not a measurement, it is a sample of one. Anyone inferring from one ChatGPT result which sources feed "the answer" is reading noise.
On Google, wording is the only variable
The reverse holds on the two Google surfaces. The identical query minutes later returns 100 percent the same sources. Every point lost below that is therefore down to wording alone: 41 points on AI Overview, 59 points on AI Mode.
That is consistent with both surfaces sitting on top of a classic result page. Change the query and you hit a different result page. How a single user question turns into several searches in the first place is something we took apart in our piece on fanout queries.
The vendors recommended are far more stable than the sources
For most tracking goals the interesting question is not the source mix anyway. It is whether who gets recommended changes. There the picture is much calmer.
Across all four engines and all ten phrasings, Monday, ClickUp and Asana appear in every single answer, Jira in at least nine out of ten. On AI Overview, Monday is the most-mentioned vendor in 8 of 10 phrasings, on AI Mode in 7 of 10.
Even where phrasing demonstrably bites, it bites vendors less than sources: on AI Overview rewording costs 41 points of source overlap but only 18 points of vendor overlap. On AI Mode, 59 against 33.
One exception is instructive. Smartsheet appears in 10 of 10 answers on AI Overview and AI Mode and in 9 on Gemini, but in only 2 on ChatGPT. So the stable core is not the same core everywhere. Within an engine it holds across all phrasings; between engines it does not.
How many of the ten phrasings named each vendor:
| Vendor | ChatGPT | Gemini | AI Overview | AI Mode |
|---|---|---|---|---|
| Monday | 10 | 10 | 10 | 10 |
| ClickUp | 10 | 10 | 10 | 10 |
| Asana | 10 | 10 | 10 | 10 |
| Jira | 9 | 10 | 10 | 9 |
| Smartsheet | 2 | 9 | 10 | 10 |
| Trello | 7 | 10 | 9 | 7 |
| Notion | 9 | 9 | – | 5 |
The columns say more than the rows. On AI Overview the list is short and closed, while on ChatGPT one vendor the other three carry throughout, Smartsheet, is almost entirely missing.
Core and long tail: which sources and vendors keep coming back
The same stability can be read from the other side: how many distinct domains and vendors show up across the ten phrasings at all, and how many of those are regulars rather than one-offs?
| Engine | Domains total | in ≥8/10 | once only | Vendors total | in ≥8/10 | once only |
|---|---|---|---|---|---|---|
| ChatGPT | 16 | 0 | 8 | 21 | 5 | 8 |
| Gemini | 26 | 5 | 14 | 25 | 9 | 5 |
| AI Overview | 14 | 5 | 6 | 9 | 6 | 0 |
| AI Mode | 43 | 6 | 24 | 12 | 5 | 2 |
Two rows stand out. On ChatGPT there is not a single domain that appears in at least eight of ten answers. There simply is no source core to track. And AI Mode draws by far the widest source base at 43 domains, 24 of which appear exactly once. Looking for stable sources there means looking in a very long tail.
On vendors the picture inverts: AI Overview names only 9 distinct vendors in total, and none of them appears just once. You will not find a more closed recommendation list than that.
The four engines recommend the same vendors but cite different sources
The comparison between engines comes out far more decisively than any phrasing effect. Of 65 source domains across the whole experiment, exactly one appeared in all four engines: reddit.com. 21 domains were cited by AI Mode alone. ChatGPT and Gemini overlapped by 3 percent on an identical question.
On vendors the same calculation lands at 38 to 68 percent. The four systems largely recommend the same companies and justify it with almost entirely different sources.

The diagonal is what makes the comparison readable: ChatGPT agrees with itself 12 percent of the time, while AI Overview and AI Mode agree with each other 42 percent of the time. Two different Google surfaces are more aligned on sources than ChatGPT is with itself.
The practical consequence: a source strategy that works for ChatGPT does not automatically work for AI Mode. And Search Console figures only cover the Google slice of that, as we described in our post on AI search visibility across Google and Bing.
Which rewording actually costs you
Since the phrasing effect only exists on the Google surfaces, the breakdown is only worth doing there. On AI Overview, best project management software, top project management software, best software for managing projects and the two plain questions form a tight block at roughly 74 percent overlap with the rest. Swapping best for top costs nothing.
Two variants fall well out of that block:
best project management toolat 43 percentrecommend the best project management softwareat 33 percent
Changing the head noun from software to tool hits a noticeably different result page on Google, even though the intent is identical. So does the imperative. Synonyms in the adjective are cheap; synonyms in the noun are expensive.
Broken down by tier, each cell being the source overlap between variants of the tiers involved:
| Tiers compared | Pairs | AI Overview | AI Mode |
|---|---|---|---|
| Word choice ↔ word choice | 6 | 78% | 43% |
| Sentence form ↔ sentence form | 6 | 55% | 43% |
| Conversational ↔ conversational | 1 | 63% | 50% |
| Word choice ↔ sentence form | 16 | 67% | 47% |
| Word choice ↔ conversational | 8 | 48% | 32% |
| Sentence form ↔ conversational | 8 | 43% | 35% |
The expensive moves are the two jumps into conversational phrasing. They cost 48 and 43 percent on AI Overview and 32 and 35 percent on AI Mode, in every case well below the comparisons within a single tier. So tracking only keyword-style prompts does not cover how conversational users behave.
One row deserves suspicion: conversational ↔ conversational rests on a single pair of variants, because only two of the ten phrasings fall into that tier. The 63 and 50 percent are therefore one comparison rather than an average, unlike the 6 to 16 pairs behind the other rows.
How to set up your prompt list for visibility tracking
Four rules follow from the numbers, and the third one contradicts the common advice:
- One phrasing per intent, not ten. Spend your budget on covering different search intents rather than variants of the same one. The stable vendor core appears in every variant, so extra phrasings buy almost no new information.
- Judge visibility at the vendor level, and treat sources separately and more cautiously. "Are we mentioned" is a robust signal. "Through which source" is a shaky one that needs more data points.
- On ChatGPT and Gemini you need repetition, not variation. Running the same prompt several times and aggregating beats ten phrasings once each. Anything else measures the model's own noise.
- On AI Overview and AI Mode the opposite is true. One run per phrasing is enough, and real variation pays off. Those variants should then match how your audience actually words it, above all in the noun.
In practice: your prompt list is not the same list for all four engines, and your run schedule certainly is not.
The ten variants in detail
For each variant, the mean source overlap with the other nine. The two footer rows supply the yardstick: where an engine's mean sits at its own noise level, every swing in that column is measurement noise rather than a property of the phrasing.
| # | Phrasing | Tier | ChatGPT | Gemini | AI Overview | AI Mode |
|---|---|---|---|---|---|---|
| 1 | best project management software | A | 23% | 40% | 74% | 56% |
| 2 | best project management tool | A | 28% | 54% | 43% | 25% |
| 3 | top project management software | A | 3% | 48% | 74% | 32% |
| 4 | best software for managing projects | A | 16% | 50% | 74% | 56% |
| 5 | what is the best project management software? | B | 29% | 50% | 74% | 56% |
| 6 | which project management software is best? | B | 29% | 46% | 74% | 56% |
| 7 | recommend the best project management software | B | 15% | 38% | 33% | 27% |
| 8 | what project management software should I use? | B | 14% | 57% | 49% | 33% |
| 9 | I'm looking for the best … what do you recommend? | C | 4% | 42% | 49% | 41% |
| 10 | what's the best … to use these days? | C | 17% | 44% | 45% | 29% |
| mean for this engine | 18% | 47% | 59% | 41% | ||
| of which own noise | 12% | 45% | 100% | 100% |
Bold marks the variants that scatter clearly more than is typical for that engine. They occur only in the two Google columns. Nothing is marked in the ChatGPT and Gemini columns, because there the mean sits essentially on the noise level and outliers cannot be told apart from randomness at all.
What this experiment does not show
Publishing numbers without their limits is not something we are willing to do, so here they are:
- One topic, one market, one point in time. We measured a broad, listicle-saturated tech category in English. Whether a narrow B2B niche behaves the same way is an open question.
- The 100 percent on Google is short-term determinism. The control runs were about eight minutes apart. How far AI Overview and AI Mode drift across days is not measured here, and real tracking runs daily.
- ChatGPT cites little. An average of 3.5 domains per answer, against 12.8 on AI Mode. Overlap values jump around on sets that small. The comparison against its own noise floor stays valid, because both sides are computed on equally small sets, but the absolute percentage should not be over-read.
- The control group is small. Three repeats per engine. That is enough to tell 12 percent from 100 percent, not enough for a precise point estimate.
The obvious next step is the same setup on a narrow non-English niche, with this run as the contrast case.
Until then the short version of the four rules holds: breadth across intents, repetition on the chat engines, and settle whether you get mentioned before asking where it came from. That is exactly what we do with B2B companies, as a GEO agency in Munich, Berlin and Zurich.