We asked Google AI Mode the same question twice, one minute apart. It agreed with itself 45% of the time.
The short answer. Across eight commercial prompts, two runs of Google AI Mode in the same minute shared only 45 to 57 percent of their cited sources. Google AI Overview, measured identically, shared 94 to 98 percent. Nothing about the world changed between the runs. The answer simply recomposed.
Why it matters. If a tool checks AI Mode once a day and charts the result, most of what you are seeing is not your visibility changing. It is the same question being answered differently.
| Surface | Agreement with itself | Reading |
|---|---|---|
| Google AI Mode | 45% to 57% | Barely better than a coin flip. |
| Google AI Overview | 94% to 98% | Effectively stable. |
8 prompts · 2 runs per surface per provider · both runs inside one minute · 31 July 2026
Method
What we ran
Eight commercial prompts, the kind a buyer types before choosing a vendor: best CRM for small business, best project management software, how to reduce SaaS churn, best accounting software for freelancers, and four more.
How we ran it
Each prompt went to Google AI Mode twice, back to back within the same minute, then to Google AI Overview twice the same way. We did this through two independent data providers so a single provider's flakiness could not masquerade as a finding, which matters because both providers showed the same pattern.
What we compared
Not the wording. Two answers can phrase the same recommendation differently and both be correct. We compared the set of source domains each answer cited, scored as the overlap between the two runs divided by their union. Identical citations score 100 percent. No shared citations at all scores zero.
The two surfaces are not the same kind of thing
This is the part worth internalising. Both are Google. Both are generative. Both cite sources. And they behave completely differently under repetition.
AI Overview behaves like a cached artifact. Ask twice, get essentially the same answer with essentially the same citations. AI Mode behaves like a fresh composition: it re-runs its retrieval and re-decides what to cite, and on our prompts it changed its mind about roughly half the sources every time.
So reporting a single “Google” number for both is not a simplification, it is an error. One of the two figures being averaged is stable and the other is close to a coin flip, and the average describes neither.
What this does to a daily chart
Consider a brand genuinely present in about a third of AI Mode answers. Checked once a day, that brand will appear some days and not others, with no change in its actual standing. Plot those checks and you get a line that rises and falls convincingly.
Someone will be asked to explain that line. They will find a reason, because there is always a reason available: a competitor published something, an algorithm update landed, the new page went live. The reason will be fiction, and it will be acted on.
That is the real cost of unsampled measurement. Not that the number is wrong, but that it is confidently wrong in a way that generates work.
What we do about it
We run every prompt multiple times per engine and pool the results, report the rate with the number of answers it came from, and attach a confidence interval to it. When a rate moves but the intervals still overlap, the product says trending rather than claiming a win.
This study is the reason that design exists rather than a preference. At 45 percent self-agreement, a single AI Mode check is not a small measurement. It is closer to an anecdote.
How to check this yourself
Open Google AI Mode, ask a commercial question in your category, and write down every source it cites. Wait a minute. Ask the identical question again and write the sources down again. Compare the two lists.
Then ask whichever tool you currently pay how many times it ran that prompt before it drew you a chart.
- Run date:
- 31 July 2026. Both runs of each prompt inside the same minute.
- Sample:
- 8 commercial prompts, 2 surfaces, 2 independent data providers, 2 runs each. 64 answers.
- Metric:
- Jaccard overlap of the cited source domains between the two runs. Wording was not compared.
- Both providers agreed:
- The volatility showed up independently in both, which is why we read it as Google recomposing rather than one vendor being unreliable.
- What this does not establish:
- How AI Mode behaves in other categories, other locales, or over longer gaps than a minute. Eight prompts on one day is enough to show the instability is real and large. It is not enough to put a precise figure on it, so treat 45 to 57 percent as the range we observed rather than a constant.
The same run also settled which data provider we use for Google surfaces, by reading the live SERP in a browser and counting. That, and the rest of how we measure, is on our methodology page. To see these rates on your own buyer questions, run a free report.
