The verdict is stable. The sources are not.
The short answer. Across 5,356 production AI answers, whether an engine names a brand holds across 90 to 99 percent of repeated runs of the same prompt. Which sources the engine cites holds far less: from 82 percent on Perplexity down to 27 percent on Google AI Mode. The two kinds of claim need completely different sample sizes, and almost every tool in this category prices and reports them as if they were the same thing.
Why it matters. A brand-visibility number from a few runs is defensible. A source-level claim from the same few runs is noise on most engines, and a dashboard that renders both with equal confidence is asking you to act on the second as if it were the first.
Corpus: 5,356 ok answers · 8 engines · 421+ prompts · 25 brands · 11 July to 27 August 2026 · full method and limits below
Finding 1: run the same prompt twice and the verdict holds, the evidence does not
For every prompt measured more than once on the same engine in the same period, we compared the runs two ways: did the named-or-absent verdict flip, and how much did the cited-source sets overlap (Jaccard).
| Engine | Verdict flipped | Source overlap | Reading |
|---|---|---|---|
| Perplexity | 1 of 137 cells (0.7%) | 82% | Repeats itself faithfully. Source-level reads are meaningful here at modest n. |
| ChatGPT (consumer) | 2 of 86 cells (2.3%) | 55% | Stable verdict; 17% of run pairs shared no source at all. |
| ChatGPT (API) | 4 of 50 cells (8.0%) | 54% | A different product from the consumer surface, measured separately. |
| Google AI Overview | 11 of 127 cells (8.7%) | 45% | Middling on both axes. |
| Gemini (API) | 5 of 50 cells (10.0%) | 32% | Two thirds of sources change between runs. |
| Google AI Mode | 12 of 136 cells (8.8%) | 27% | Recomposes almost everything except the verdict. |
multi-run cells per engine as shown · 3,392 source run-pairs total · pairs where either run cited nothing are excluded from overlap (vacuous agreement is not stability)
The spread across engines is more than 3x, which has a practical consequence: one stability threshold applied to every engine mislabels most of them. Our product carries a measured per-engine baseline instead, and renders a movement smaller than the engine’s own self-disagreement as within normal variation.
Finding 2: the engines barely agree with each other, including two from the same company
Same prompt, same period, different engines, compared by the union of sources each engine cited across its runs.
| Pair | Comparisons | Mean overlap |
|---|---|---|
| ChatGPT API x ChatGPT consumerThe best pair in the whole corpus is the same vendor's two surfaces, and they still agree on only a third. | 71 | 34% |
| Google AI Mode x Google AI OverviewOne company, one 'seamless experience' per its marketing, one shared source in five. | 472 | 20% |
| Copilot x Google AI ModeNearly half of these comparisons shared nothing at all. | 272 | 5% |
No cross-vendor pair clears 25 percent. Averaging engines into one visibility score therefore averages away the only diagnostic signal in the data: which engine is missing which evidence. Engine disagreement is not noise. It is the pointer at what to change.
Finding 3: the number nobody publishes, the effective sample size
Answers to the same prompt correlate: a prompt the engine answers with your brand today mostly gets your brand tomorrow too. The sampling literature calls this clustering, and it means the nominal answer count overstates the information collected. We measured the intraclass correlation of the named verdict across prompts, per engine, and computed what the nominal sample is actually worth.
| Engine | Nominal answers | Prompts | ICC | Effective n |
|---|---|---|---|---|
| Perplexity | 860 | 235 | 0.92 | 250 |
| Google AI Mode | 773 | 216 | 0.75 | 263 |
| Google AI Overview | 678 | 211 | 0.83 | 238 |
| ChatGPT (consumer) | 606 | 117 | 0.80 | 139 |
one-way ANOVA ICC over prompts with 2+ runs · design effect 1 + (mean runs per prompt - 1) x ICC
Read the consumer ChatGPT row: 606 answers were collected, and after clustering they carry the information of about 139 independent ones. With correlations this high, the tenth repeat of a prompt you already measured buys almost nothing, while a new prompt buys close to a full unit of information. If a vendor sells you more runs per day on the same prompts, this table is what they are selling against. Diversity of prompts is the axis that pays.
Finding 4: on some surfaces, most answers cite nothing, and that is not absence
The share of successful answers that cited no source at all, by surface. Any citation metric that counts these as “you were not cited” is manufacturing zeros:
- ChatGPT (API): 339 of 533 (63.6%)
- ChatGPT (consumer): 367 of 869 (42.2%)
- Claude (API): 30 of 282 (10.6%)
- Gemini, both Google surfaces: 1.3% to 2.3%
- Copilot, Perplexity: 0 of 274 and 0 of 1,043
The two ChatGPT rows are the same brand name and 21 points apart, which is one more reason a “ChatGPT visibility” number is meaningless until the surface is named.
Finding 5: the cited web rots, though less in our corpus than the published rate
We probe the URLs the engines cite for our tracked brands and record link health with a conservative ladder: only a definitive 404 or 410 counts as dead, only a host that no longer resolves counts as gone, and a bot gate or timeout counts toward nothing. First sweep across 1,500 cited URLs: 17 dead of 1,289 decided (1.3 percent), with 135 bot-gated and 76 unreachable excluded from the claim. The published category figure is 19.3 percent dead; our corpus skews toward B2B pages, so both numbers can be right about different webs. Either way, among the 17 were two pages from a tracked competitor that engines still cite, which is the actionable version of this finding: a dead cited page is a vacant slot with proven demand.
Method
The corpus is every successful brand-level answer our production measurement captured from 11 July to 27 August 2026: 5,356 answers across 8 engines and 421+ prompts for 25 brands. Presence verdicts come from a word-boundary matcher with typography normalization; we validated it three ways before trusting any number here: an adversarial characterization set with published per-family results, a two-matcher cross-check over 6,001 answers that found zero disagreements, and a stratified 30-sample hand check that read 30 of 30 correct. Source overlap is Jaccard over cited domains. Consumer surfaces are scraped as a user sees them; API surfaces are measured separately and labeled, never blended.
Limits, stated rather than buried
- 25 brands in a handful of mostly B2B verticals. This is a real measurement of a narrow slice, not a web-representative sample.
- The corpus grew around one large July baseline, so today’s rerun agreeing with the July numbers is consistency on extended data, not an independent replication.
- Copilot and Claude have too few repeated-run cells to publish stability numbers; they render as unmeasured in the product rather than borrowing a neighbour’s figure.
- The link-health rate is one sweep over one corpus. It will move as coverage widens, and the page will be updated with the date attached.
What this means if you are buying measurement
Ask any vendor three questions. How many runs per prompt per engine per day, stated as a number? Which claims does that sample support, brand-level or source-level? And what is the effective sample after clustering? The IAB’s August 2026 standard now requires disclosing most of this; our standing against it, including the criteria we only partially meet, is published here.
Or measure it yourself: the first baseline is free and every number in it carries its denominator.
