Research · Original measurement · 2 September 2026

How eight AI engines answer the same questions

The short answer. Asked the same buyer questions, the engines do not behave like one thing called AI search. Perplexity cites on every answer and repeats 89% of its sources when asked again; Google AI Mode cites almost as often but keeps only 24% of its sources between two runs of the same prompt. The consumer ChatGPT cites on 65% of answers, the ChatGPT API on 24%. The best agreement between any two engines on the same prompt on the same day is one source in five.

Why it matters. A single blended visibility score averages engines that share almost no sources, and a source-level claim on some engines needs many more answers than a brand-level claim. Which engine, which surface, and how many answers are behind a number are not footnotes. They are the number.

Corpus: 10,639 valid answers · unbranded prompts only · brand-level runs · 9 engine surfaces · 11 July to 2 September 2026 · 77,526 citations over 14,477 domains · per-engine CSV · cross-engine CSV · long-tail CSV

Finding 1: the engines are nine different instruments

One row per engine surface. “Cited a source” is the share of valid answers that carried at least one linked source. “Sources per answer” is the mean over valid answers. “Source overlap” is the Jaccard overlap of cited domains between two answers to the same prompt on the same day, averaged over every such pair. “Verdict flip” is how often the brand-named verdict differed between those two answers, over pairs where both answers were graded.

EngineAnswersCited a sourceSources per answerReddit shareRepeat pairsSource overlapVerdict flip
ChatGPT (consumer)30 Jul to 2 Sep2,86365.0%3.81.44%1,89933.7%7.10%
Perplexity11 Jul to 2 Sep2,224100.0%9.57.25%2,07888.7%2.29%
Google AI Overviews11 Jul to 2 Sep1,88392.3%9.04.82%1,56145.8%3.36%
Google AI Mode11 Jul to 2 Sep1,33695.4%12.43.11%1,28223.9%2.45%
Copilot3 Aug to 2 Sep727100.0%3.50.16%42565.1%4.03%
Gemini (consumer)26 Aug to 2 Sep51485.4%3.23.31%42535.7%2.08%
ChatGPT (API)11 Jul to 3 Aug42923.8%1.817.81%11724.2%6.84%
Gemini (API)11 Jul to 3 Aug42697.2%12.71.21%11524.4%8.70%
Claude (API)16 Jul to 5 Aug23788.6%6.60.00%n/an/an/a

Two things stand out. Source stability spans a factor of almost four, from 24% on Google AI Mode to 89% on Perplexity, so one repeat-run rule for every engine mislabels most of them. And Google AI Overviews failed 8.0% of probes (no overview was returned where one was expected) against 0 to 0.7% everywhere else, which is why its coverage is thinner than its answer count suggests and why the product says so on its own page.

Finding 2: a surface is not its API

The consumer ChatGPT cited a source on 65.0% of answers, with 3.8 sources per cited answer on average. The ChatGPT API, asked the same questions over its own window, cited on 23.8% with 1.8 sources. Their Reddit shares run the other way: 1.44% of the consumer surface’s citations point at reddit.com against 17.81% of the API’s. Gemini shows the same split in the other direction on citation rate (97.2% on the API, 85.4% on the consumer app). A dashboard that calls both things “ChatGPT” is reporting two different products under one name.

Finding 3: engines barely agree with each other

For every prompt measured on two engines on the same day, we compared the sets of domains each cited. The eight best-agreeing pairs, over at least 100 comparisons each:

Engine pairSame prompt, same dayShared sources
Gemini (consumer) and Google AI Overviews16218.0%
Google AI Mode and Google AI Overviews55717.7%
Google AI Overviews and Perplexity1,02716.2%
Gemini (API) and Google AI Overviews30913.6%
Gemini (API) and Perplexity33713.3%
Claude (API) and Perplexity20812.0%
Google AI Mode and Perplexity63711.5%
Gemini (consumer) and Perplexity19511.5%

The best pair in the corpus is two Google surfaces sharing 18% of their sources. Blending engines into one score does not average out noise; it averages away the only signal there is, which is that each engine trusts its own sources. A fix aimed at one engine’s sources should be expected to leave the others where they were.

Finding 4: citation is a long tail, not a shortlist

The 10,639 answers carried 77,526 citations across 14,477 distinct domains. The ten most-cited domains account for 15.6% of citations, the top fifty for 26.8%, the top two hundred for 40.1%. Six citations in ten point somewhere outside the two hundred most-cited domains. “Get on the handful of sites the engines trust” describes a minority of what the engines actually cite.

Method

Each prompt is asked on each engine on a schedule, usually several times per day of measurement. Every answer is stored whole with its linked sources. A run that returned an error or no surface is a failed probe and is excluded from every rate above but counted in the failure column. Domains are lowercased with a leading www removed; Google redirect wrappers are resolved to the real cited site before counting (a repair dated 28 August in our measurement changelog). The brand verdict on an answer is the graded stance; answers not yet graded are excluded from the flip column and reported in the CSV as the gap between repeat pairs and graded pairs.

Reconciliation with earlier figures

An internal cut of this corpus dated 26 August (6,112 answers) put the consumer ChatGPT citation rate at 57.1% and its run-to-run source overlap at 54.5%; the figures here are 65.0% and 33.7%. Two things changed: the corpus nearly doubled, and the pair definition was tightened to consecutive answers within the same prompt, engine and day, the definition the base-rate study uses. Google AI Mode moved from 26.9% to 23.9% and Perplexity from 82.4% to 88.7% under the same change. The ordering of engines by source stability did not change. Earlier figures stay in the record; a correction is a new dated entry, never an edit.

What this does not prove

  • It is not a random sample of the web. This is an observational production cohort of our customers' buyer questions, weighted toward B2B software and services, in English, measured from one region.
  • A citation rate is not a quality score. An engine that cites on every answer is not thereby more accurate, and an engine that cites rarely is not thereby worse; the two ChatGPT surfaces show the same model family behaving differently by product design.
  • Verdict flip rates are read against each engine's own base rate of naming a brand, not against each other. The null model and the correction are in the companion study linked below; the raw flip figures here are the inputs to it, not a ranking.
  • Cross-engine overlap is measured on the same prompt on the same day, which is the fairest comparison we can make, but the engines were not asked at the same minute and each was asked in its own product shape.
  • The three retired API surfaces ran for shorter windows on fewer prompts. Their rows are kept because the surface-versus-API gap is a finding, not because their numbers are as settled as the consumer rows.
  • Nothing here is causal. No content was changed to produce these figures, and no engine was asked twice in a way that could influence its second answer.

Related: the noise floor each engine sets for its own movement, in The noise floor of every engine.