The ChatGPT you measure is not the ChatGPT your buyers use.
The short answer. We asked the same commercial questions of the OpenAI API and of chatgpt.com within the same minute. The two surfaces recommended almost entirely different providers. On the first prompt they shared one domain out of seven. On the second they shared only Trustpilot and Reddit, and not a single recommended service. A brand that measures ChatGPT through the API alone is not measuring what its buyers see.
The limit, stated up front. This is three prompts in one category on one day. It is enough to prove the two surfaces can diverge, which is what we needed to decide our own roadmap. It is not enough to tell you how often they diverge, and we do not report a rate, because three prompts cannot support one.
Almost every tool that claims to track ChatGPT is querying the OpenAI API. It is the reasonable engineering choice: it is documented, it is stable, and it returns clean structured output. We built our own ChatGPT adapter that way too.
The assumption underneath it is that the API answers roughly like the product. We wanted to know whether that assumption survives contact with evidence, because if it does not, then a whole category of tools is confidently reporting the wrong thing.
Method
What we compared
Two surfaces of the same brand. The API surface is the OpenAI API plus its hosted web search tool, which is what most visibility tools mean when they say they track ChatGPT. The consumer surface is chatgpt.com itself, captured through DataForSEO's LLM scraper, which returns the real answer markdown with its inline citations, the fan-out queries behind it, and any sponsored items.
How we ran it
Fresh on both sides, within the same minute, on three commercial prompts in one category. We first tried comparing fresh consumer answers against stored API measurements and abandoned that design: no brand in our database had rich enough stored evidence to make the comparison fair, and a comparison against a weaker record would have flattered the finding.
What we counted
The domains each answer cited, and the products each answer actually recommended. Those are different things, and the distinction turned out to carry the result.
What the three runs returned
“best app to create a US passport photo”
- smartphone-id.com
- photoaistudio.com
- itseasy.com
- snap2pass.com
- snap2pass.com
- play.google.com
- pixid.studio
- apps.apple.com
Overlap: 1 domain of 7 unique. One brand appeared on both sides (Smartphone iD). The consumer answer carried one ad.
“best online passport photo services”
- 8 domains
- compliantphoto.com
- pixid.studio
- photoaistudio.com
- 6 domains
- passport-photo.online
- epassportphoto.com
- snap2pass.com
Overlap: 2 domains, both generic. The only shared sources were trustpilot.com and reddit.com. Every recommended service was different. The consumer answer carried one ad.
“best PhotoAid alternatives”
- 5 domains
- empty response
Overlap: not comparable. The consumer scrape returned nothing. We report it as a failed run, not as an answer that named nobody.
The second prompt is the one that matters
Counting shared domains understates the gap. On the second prompt the two answers did share two sources, so a naive overlap metric would score it as partial agreement. But both shared sources were Trustpilot and Reddit, which are review aggregators rather than providers. Of the businesses actually recommended to a buyer, the overlap was zero.
That is the distinction a visibility tool has to get right. Being cited as a source and being recommended as the answer are separate outcomes, and a brand can win one while losing the other.
It is why we never collapse them into a single visibility score. Presence is measured as three separate rates for every prompt: how often an answer names you, how often it cites one of your own domains, and how often it does both in the same answer. Then a second layer grades how the answer actually presented you, and that grade only counts when the classifier can produce a verbatim sentence from the answer that names you. If it cannot, the run stays unevaluated and shrinks the sample rather than quietly counting as “not recommended”.
Two things we found that we were not looking for
Ads are already inside consumer answers. Two of the three consumer responses carried a sponsored item. Nothing in the API surface shows them. If you are benchmarking your category through an API, paid placement in the answer your buyer reads is invisible to you.
The consumer surface can return nothing. One of three scrapes came back empty. We record that as a failed run and exclude it from every denominator. The tempting alternative, counting it as an answer that mentioned nobody, would quietly depress the measured rate of every brand in the category and make our own product look more useful than it is.
What we changed because of it
We ship the consumer surface as its own engine rather than folding it into our ChatGPT number. Averaging the two would have produced a figure that describes neither. They appear as separate rows, measured separately, and a brand can see that it is winning one and losing the other.
We also wrote the finding into our published refusals: we will not sell AI answer measurement drawn from model APIs alone. That page lists the other things we decline to build and why, including our position that schema markup is hygiene rather than a citation lever.
How to check this yourself
You do not need our product to reproduce the shape of this. Ask ChatGPT the question your best customer would ask, in the app, and write down every business it names. Then ask the same question through the API, or through any tool that says it tracks ChatGPT, and compare the two lists. If they differ, the number your tool reports is describing a surface your buyers do not use.
If you would rather have it measured repeatedly and with sample sizes attached, that is what we built. A free report runs your own buyer questions across the engines and shows the rates with their ranges, including the runs that failed.
- Run date:
- 30 July 2026, both surfaces within the same minute.
- Sample:
- 3 commercial prompts, 1 category (passport photo services), 1 run per surface per prompt.
- Failed runs:
- 1 of 3 on the consumer surface, reported and excluded, never counted as a zero.
- What this does not establish:
- how frequently the surfaces diverge, whether the gap holds in other categories, or any figure that would need a larger sample to be honest. We report no rate from this study.
More on how we measure, and the tactics we refuse, on our methodology page.
