What we measured, including the parts that weaken the finding.
We publish our own research the way we report a customer’s numbers: with the method, the sample size, and the runs that failed. Where a study is too small to support a rate, we say so and report no rate.
These are the field notes. The formal studies, each with a frozen cohort and a downloadable aggregate, are in Research.
The verdict is stable. The sources are not.
5,356 production answers across 8 engines. Whether a brand is named holds across 90 to 99 percent of repeated runs; which sources the engine cites holds as little as 27 percent. Plus the number nobody publishes: the effective sample size after clustering, and what it says about paying for more runs.
Google AI Mode agreed with itself 45% of the time
Same eight prompts, twice, one minute apart. AI Mode shared only 45 to 57 percent of its cited sources with itself. AI Overview, measured identically, shared 94 to 98 percent. Nothing changed between the runs except the answer.
The ChatGPT you measure is not the ChatGPT your buyers use
Same commercial prompts, same minute, two surfaces. One shared domain out of seven on the first prompt, and zero shared recommendations on the second. Full per-prompt data, including the run that came back empty.
We fact-checked a Google AI Mode answer against the vendor's own docs
An AI Mode answer described a data vendor's product as covering assistants its own documentation does not commit to. We were the buyer, and the gap changed our integration plan.
