How to measure your brand's visibility in AI answers

The short version. Pick the unbranded questions your buyers actually ask, run each one several times across several engines, and record the share of answers that name you along with the number of runs. A single check on a single day is an anecdote, not a measurement.

What you need

  • A list of the questions a buyer asks before choosing a vendor in your category.
  • Access to the engines you care about, either directly or through a tool.
  • A spreadsheet, if you are doing this by hand.
  1. 01Write prompts a buyer would actually type, without your name in them

    Prompts containing your brand name are the least informative measurement you can take. An engine asked about you will nearly always discuss you, so the result tells you almost nothing about whether you get discovered.

    The prompts that matter are unbranded and commercial: the ones where the engine chooses who to name. Phrase them the way a person talks to an assistant rather than the way someone types into a search box, because the phrasing changes the answer.

  2. 02Decide what counts as a win before you look

    Fix your definitions first, or you will move them later to suit the result. Three are worth tracking separately: whether the answer names you at all, whether it cites one of your own pages as a source, and whether it does both in the same answer.

    Keep being named separate from being recommended. An answer can name you while recommending a competitor, or name you as the expensive option. Collapsing those into one score hides the gap that loses deals.

  3. 03Run each prompt many times, on each engine

    This is the step people skip, and skipping it invalidates everything after it. AI answers vary between runs, so a single response is one draw from a distribution.

    Run each prompt repeatedly on each engine and record every result. Treat each engine separately rather than averaging them: they cite different sources and reporting one blended figure describes no engine accurately.

    Measure the consumer app surface separately from the API where you can. Our own same-minute comparison found the two recommending almost entirely different providers on the same commercial prompts.

  4. 04Record failures as failures

    Engines time out and return empty responses. When that happens, exclude the run from your denominator entirely. Do not count it as an answer that failed to mention you.

    This matters more than it sounds. Counting failures as absences depresses every rate, and it makes the problem look worse than it is in a way that conveniently favors whoever is selling the fix.

  5. 05Report the rate with its sample size, and never a position

    Write the result as a share with its denominator: named in 14 of 60 answers, not 23 percent alone. The denominator is what makes the number checkable.

    Do not report a rank for AI answers, averaged or otherwise. Position metrics of that kind are typically computed only over the answers where the brand already appeared, which means the metric improves as visibility falls: the few answers still mentioning you tend to mention you prominently.

  6. 06Re-measure before claiming a change

    A second measurement after the work ships is what turns activity into evidence. Compare the two rates with their sample sizes.

    If the plausible ranges around the two rates overlap, you have not demonstrated a change yet, however much you would like to have. Say trending, keep working, and measure again with more runs.

What this method does not do

  • Sampling by hand is slow and gets abandoned, which is the honest reason most teams end up buying a tool. The method above is the same either way.
  • Rates measured on one day describe that day. Engines update, and a measurement is a snapshot rather than a permanent property of your brand.

Every step above works without buying anything. If you would rather have it run repeatedly with the sample sizes attached, our free report does the measurement on your own domain.