Research · Pre-registered method · 2 September 2026

The census: a pre-registered method

What this is. A census is a fixed set of buyer questions asked on every consumer AI surface, sampled until each cell is settled, and published with its counts. This page is the method, written and dated before the first answer is collected, so the result can be held to it.

Status. Pilot pending. The pilot frame is written; collection starts when its spend is approved, and the first readout will be dated here.

The method

  1. 01 · The frame

    A fixed list of buyer questions per category, written as a buyer would type them, in several wordings each. The list is frozen and published with the results; questions cannot be added or removed after collection starts.

  2. 02 · The surfaces

    The consumer surfaces buyers use, measured as themselves: ChatGPT (chatgpt.com), Google AI Overviews, Google AI Mode, Gemini (gemini.google.com), Perplexity and Copilot. Developer APIs are not stand-ins for any of them; the engine behaviour study shows why.

  3. 03 · The unit

    A cell is one question wording on one surface. Every answer is stored whole, with its linked sources, and graded for whether each roster brand is named, cited and recommended. Named and cited are deterministic; recommended is graded and its grader is named.

  4. 04 · The stopping rule

    Cells are sampled until the 95% interval on the named rate is narrower than a stated width, or until the cell reaches its ceiling. Published sampling work puts the point at which citation rankings settle at 33 to 75 citation-bearing answers per cell, and finds that repeats past the fifth on one wording buy very little, so budget goes to wordings and languages before it goes to repeats. Every cell reports its own n.

  5. 05 · The reporting

    Counts before shares. Every rate carries its numerator, its denominator and a Wilson interval. Effective sample sizes are reported after clustering by question, because answers to the same question are not independent. Engines under eight answers on a cell render as counts.

  6. 06 · What is published

    Per-category tables by surface, the frozen question frame, the aggregate CSV, the grader and its agreement check, the failure counts per surface, and a dated reconciliation section if any figure is later corrected.

Why a stopping rule and not a fixed n

A fixed number of runs per question is either wasteful on settled cells or insufficient on contested ones. The number of answers a cell needs depends on how close its named rate sits to a coin flip and on how many distinct sources the surface cites. Sampling to an interval width spends answers where the uncertainty is, and reports the n it took, which is itself a finding: a cell that will not settle is a contested question, and contested questions are where the work is.

What the census will refuse to claim

  • No ranks. A brand's position inside an answer is not measured because it is not stable enough to be a number.
  • No causal claims. The census observes; nothing is changed on any site to produce it.
  • No claim of representativeness beyond the frame. The frame is a designed list, not a sample of what buyers ask.
  • No blended score across surfaces. Each surface is reported on its own because they cite different sources.
  • No single-run numbers. A cell that did not reach its stopping condition is reported as unsettled, not as a rate.

The reporting contract is the one every page in the product follows; the noise floors it applies are in The noise floor of every engine.