The noise floor of every engine
The short answer. Ask an engine the same question twice on the same day and it disagrees with itself. How much depends on the engine and on what you are looking at: the brand verdict (named or not) flips on 2 to 9% of repeated pairs, while the cited sources change on anything from 11% of pairs (Perplexity) to 76% (Google AI Mode). That self-disagreement is the floor a week-to-week move has to clear before anyone is allowed to call it a change.
Why it matters. Most dashboards draw a line and let the reader infer that a wobble means something. A floor measured per engine is the difference between a chart and an instrument.
Product rule measured 27 August 2026 · fresh recomputation 2 September 2026 over 7,902 repeat pairs · same prompt, same engine, same day · per-engine CSV
The rule the product applies today, beside the fresh number
“Source overlap” is the share of cited domains two answers to the same prompt have in common (Jaccard). “Verdict flip” is how often the named-or-not verdict differed between them. The product columns are the constants in the software as of this writing; the fresh columns are the same quantities recomputed on 2 September.
| Engine | Rule: source overlap | Rule: verdict flip | Fresh pairs (graded) | Fresh: source overlap | Fresh: verdict flip |
|---|---|---|---|---|---|
| Perplexity | 82.4% | 0.7% | 2,078 (1,922) | 88.7% | 2.29% |
| ChatGPT (consumer) | 54.5% | 2.3% | 1,899 (1,704) | 33.7% | 7.10% |
| Google AI Overviews | 44.9% | 8.7% | 1,561 (1,371) | 45.8% | 3.36% |
| Google AI Mode | 26.9% | 8.8% | 1,282 (1,100) | 23.9% | 2.45% |
| Copilot | not set | not set | 425 (298) | 65.1% | 4.03% |
| Gemini (consumer) | not set | not set | 425 (288) | 35.7% | 2.08% |
| ChatGPT (API) | 53.7% | 8.0% | 117 (117) | 24.2% | 6.84% |
| Gemini (API) | 31.8% | 10.0% | 115 (115) | 24.4% | 8.70% |
How the floor is used
Three places in the product read these constants. When a brand asks why a number moved, the answer states the engine’s own repeat-run disagreement beside the move, so a wobble inside the floor is named as wobble. On the sources view, a change in an engine’s cited sources is flagged as movement only when the overlap with the prior run falls more than ten points below that engine’s floor. And engines with fewer than eight answers behind a rate render as counts, not percentages, because no floor can be applied to a number that small.
Reconciliation between the two columns
The 27 August rule was computed over pairs within the same measurement period; the 2 September column uses consecutive answers to the same prompt on the same day, over a corpus nearly twice as large, and takes the verdict from the graded stance, excluding pairs where either answer is ungraded. Under that definition the consumer ChatGPT’s source overlap reads 33.7% against a rule of 54.5%, and its verdict flip 7.1% against 2.3%. The direction of every other engine held. The rule will be re-pinned to the fresh definition with a dated changelog entry, and until then the product keeps the older constants because a rule that changes without a date is not a rule.
What this does not prove
- A noise floor is a property of our measurement of an engine, not of the engine. Different prompts, a different region or a different time of day would give a different floor.
- The two columns use different pair definitions and the fresh column comes from a corpus nearly twice the size, so the gaps between them are partly definitional. The reconciliation section says which is which.
- A move that clears the floor is not thereby caused by anything the brand did. Clearing the floor is the entry condition for asking why, not the answer.
- Verdict flips are unconditional rates and depend on each engine's base rate of naming a brand. The companion study on flip rates gives the null model; read the flip column with it.
- Copilot and the consumer Gemini surface have no product rule yet because their windows are short. The product treats them as if the floor were unknown, which is the conservative reading.
Related: How eight AI engines answer the same questions and A flip rate is not a stability metric.
