Methodology
Receipts, not vibes: how we measure.
AI answers are noisy, and most visibility tools report that noise as fact. Ansengine is built on 6 rules that are enforced in code, not in marketing. Last updated 20 September 2026.
Enforced in code is easy to say, so the enforcement page names the file and the test behind each rule, and marks which ones you can falsify yourself from outside.
1. Statistics, not point estimates
Every rate ships its denominator and a Wilson 95% confidence interval. A prompt is measured over repeated runs, not a single screenshot.
A lift is claimed only when the interval measured after the fix sits entirely above the baseline interval at the engine's propagation window (6 hours for Perplexity up to 7 days for ChatGPT and Claude). Anything else stays labeled as measuring or directional. We never report rank positions as stable facts, because across repeated runs they are not. The rate a receipt is judged on is the cited rate: the answer names you and cites one of your own domains. Work that earns a mention without a link cannot confirm a receipt here, and the receipt says which rate it read.
What that interval does not do, stated rather than buried. A pooled rate sums counts across prompts, days and engines and treats every answer as an independent draw. Answers to the same question are not independent, and our own study of 5,356 production answers measured design effects of 2.8 to 4.4 on the same corpus, which means a pooled interval here is narrower than the data has earned. We publish that measurement instead of quietly widening the band, and the interval that accounts for clustering is on the Planned list at the foot of this page.
2. Verified facts, not model narratives
What engines assert about your brand is extracted as verbatim sentences and put in front of you for a Correct or Incorrect verdict. Verified claims become the only ground truth our agent and our drafts may use; incorrect claims become correction targets.
We never publish a model-written brand perception or head-to-head verdict. Comparisons on our surfaces are computed from measured rates only. This matters because entity confusion is real: monitoring tools have been observed describing a different company's product as the brand under measurement, with a confidence score attached.
3. Identity: aliases and entity grounding
Mention detection is word-boundary and case-insensitive across the brand name and its declared aliases, so spelling variants count and near-names do not.
The strongest rate we report, supported, requires the answer to name the brand AND cite one of the brand's own domains in the same answer. That is entity-grounded by construction.
4. Measurement hygiene
The headline score is computed on unbranded prompts. A prompt that names your brand guarantees a mention, so blending it into visibility flatters everyone; we report branded rates separately, never mixed in.
Measurement runs in the market country you declare on the brand (45 supported countries, each addressed by the provider's own location table), and your account's IP is never inherited into it. A brand with no declared market measures as a United States search, stated as the default rather than left implicit. Every geo-scoped read follows the same setting since 7 September 2026: the answer scrapes, demand volume, domain authority, daily positions, business profile and reviews, and the mentions index. Rows stored before that date for a brand outside the United States were measured as a United States search and are kept unchanged; the changelog entry says so.
When you set a market on a prompt, we localize the prompt TEXT by naming that market in it, either by filling a {city} placeholder or by appending it. That rewrite is deliberate and it is visible: since 7 September 2026 the exact string we sent is stored on every answer and shown on the receipt, so you can see the prompt as asked rather than the prompt as typed. Answers captured before that date carry no sent string and the receipt says so, because reconstructing it later would be a fabricated record.
The competitor roster is declared and editable, with rivals discovered in answers surfaced for you to confirm or reject. A wrong auto-discovered peer set can silently poison every downstream number; ours cannot define your report without you seeing it.
5. Receipts for everything
Every fix, publish, and outcome writes an auditable record: drafted placements carry an immutable draft hash through approval, CMS exports record what was pushed where in draft or live mode, and prove outcomes store both confidence intervals with the verdict.
The shared client report renders from these records at request time. Sample data in the product is always labeled Sample and can never be mistaken for a measured result.
6. Slot regimes: the right lever per pool
For each buyer query we classify the answer slot by what actually holds it in the measured answers, not by guesswork. Over its runs, the share of sourced answers that leaned on an aggregator (a directory or a best-of listicle) puts the pool in one of five regimes: sparse, mixed, directory, listicle, or knowledge. Each maps to one lever: an own-site entity page, directory presence, listicle placement, or brand mentions.
The classification is deterministic and spends nothing: it reads stored answers only. Aggregator hosts are never reported as competitors you could beat, and platform hosts such as YouTube are labeled as platforms, not rival brands. A pool is called open only when no measured answer for it has ever recommended you, and taken only when a pool that was never recommended before becomes recommended in a later period. Both are measured facts, never asserted.
The sampling dispute, reconciled
The category is publicly split on how often to sample AI answers. The split dissolves once aggregates and single prompts are treated as the different statistical objects they are, and the resolution is checkable from any vendor’s own published quotas.
Both published positions are right, about different questions
One camp has published that daily single checks are enough because brand-level results are stable day to day. Another has published that five repetitions of the same prompt still leave a margin of error near 27 points. These sound contradictory. They are not: the first statement is about AGGREGATES and the second is about SINGLE PROMPTS, and the arithmetic of sampling treats them completely differently.
A brand-level number pools hundreds of answer samples: 100 prompts on 3 engines over a month is thousands of draws, and a pooled interval narrows with the square root of that total. One prompt checked once a day is a sample of one from a noisy distribution, and its daily reading carries an interval spanning tens of points. So a daily aggregate line is honest at one check per prompt per day, and a daily PER-PROMPT verdict at the same cadence is a coin report wearing a percent sign.
The quota arithmetic you can run on any pricing page
Take a tool that advertises daily sampled measurement and publishes its monthly answer quotas. Divide the quota by prompts, engines, geos, and days. On one leading tool's published tiers the division comes out to exactly 1.0 in every configuration: fifty prompts on three engines in one geo with a 4,500-answer month is one sample per prompt per engine per day, and the same identity holds on their larger tiers. Sampled, in that catalog, means a sample size of one.
We publish the opposite arrangement instead of hiding the same one. Every prompt runs weekly on every engine in the plan, pooled numbers carry the pooled n, and a per-prompt rate is not rendered as a percentage until at least 8 checks have accumulated; below that you see the counts themselves. The interval is printed beside every rate, which is the disclosure that makes any cadence honest.
Where the daily dollar goes: the tripwire
Checking every prompt every day mostly re-measures prompts that have not changed. The variance of a yes-or-no rate peaks where the rate is contested and collapses at the stable extremes, so a daily check buys the most detection power on the contested prompts and almost none on the settled ones. Allocating daily checks by that variance reaches the same pooled precision as uniform daily checking with roughly 43 percent fewer runs.
So paid plans run a daily tripwire on the contested quarter of their prompt pool, on the highest-traffic consumer surface. A detected flip opens a burst window that re-checks the prompt daily for a week, which is what buys the follow-up sample a movement claim actually needs. The daily story is adaptive and automatic; nobody pays for daily-everything they will never read.
The stopping rule: how many runs each prompt gets
Every measurement cycle plans its runs per prompt per engine from that cell's own recent history, under one hard rule: the plan never exceeds what a flat allocation would spend. A cell whose verdict is settled (clearly absent, or clearly named with the interval to back it) drops to a single sentinel probe; the runs it no longer burns go to the cells where the verdict is genuinely uncertain, closest to a coin flip first, because that is where a run buys the most information. A cell with too little history keeps the default until it has earned a classification, and a detected flip forces a cell back to full attention whatever its history said.
What that rule buys today, stated exactly, because the cadence changes the answer. Every scheduled sweep asks for one run per prompt per engine: the weekly sweep, the daily tripwire, the priority lane and the sentinel lane all pass one. At one run the flat allocation is also the floor, so the planner hands back the flat plan and reallocates nothing. On a cycle that buys two or three runs a cell, which today is the burst that follows a detected flip, runs do move to the contested cells. On a cycle that buys four runs a cell or more, the onboarding baseline and the paid snapshot, the settled cells still drop to one probe and that money is saved rather than moved. So read the rule as the spending ceiling it enforces on every cycle, not as a reallocation you are getting on the weekly sweep. Reallocation on the weekly cadence is on the Planned list at the foot of this page.
The sentinel is why settling a prompt never blinds us to it: one probe per cycle still notices a change, and the burst machinery then escalates it to daily checks again. Real run counts land in every snapshot and in the raw exports exactly as spent, so every published rate keeps its true n. We did not find a runs per prompt figure published by any of the products on our compare pages.
Question archetypes: one headline hides a mix
The largest published archetype study (22,295 answers, 115,843 citations across 460 B2B prompts) found that the KIND of question alone moves a brand's mention rate by 8 to 17 points: comparison and recommendation asks beat research asks by 6 to 9 points on every engine measured. A single blended rate silently averages those populations, and a prompt panel whose mix drifts can move the headline without a single answer changing.
So the console reports the named rate per archetype (best-of, alternative, how-to, category), each with its own interval and run count, and discloses the panel mix beside them. When the measured mix moves materially inside a reporting window, the dashboard says so in plain terms, with the run counts, because a trend across a panel change is not a like-for-like comparison. Holding the mix constant is part of what makes a movement claim honest.
ChatGPT is two engines: the reasoning-mode disclosure
A controlled study on GPT-5-era ChatGPT (100 prompts, run in both modes) found the Instant and Thinking modes share only 25.6 percent of their cited domains. Thinking mode cited sources in 68 percent of answers against 50, cited 4.5 sources against 2.6, issued over four times the retrieval queries, leaned 363 percent harder on government and academic sources, and cited Reddit half as often. Same product name, materially different retrieval behavior.
Our chatgpt_web engine measures the DEFAULT consumer surface, which is what the large majority of buyers receive, and we say that here rather than implying coverage of every mode. Where a reasoning mode matters to a programme, the honest treatment is a separately labeled surface with its own rates, never a blend, exactly as we refuse to blend developer APIs into consumer numbers today.
Published by them, not measured by us
“Peec AI executes each prompt once every 24 hours on every AI model you’ve selected.” And: “If you are running 25 prompts across 3 models for 30 days, we would analyze 25x3x30 = 2250 AI answers.”
Growth plan: “100 unique prompts / 9,000 responses monthly” across 3 answer engines. 100 x 3 x 30 = 9,000.
“A response is a single AI answer to one of your tracked prompts on one engine. We count every fresh answer captured during a refresh, so a prompt tracked across 4 engines with daily refresh uses 4 responses per day.”
All three formulas are one answer per prompt per engine per day: a sample size of one behind every daily per-prompt reading.
Where we stand against the IAB standard
On 3 August 2026 the IAB published Measuring Visibility in the AI Era, developed with Walmart, Microsoft Clarity, Acxiom and WPP Media. It is the first industry standard for this problem, and its central sentence matches what this page has said since before it existed: “Single-response measurement is not measurement. A brand’s visibility on a given query is a distribution, not a value.” Below is our standing against its directional-versus-decision-grade criteria, stated per criterion with the gaps left as gaps. A disclosure that only ever says “meets” is marketing.
Query volume
partial, statedThe standard: A measurement programme under 50 queries is exploratory, below even directional.
Plan pools run 40 to 600 tracked prompts. Our smallest tier tracks 40, which the standard classes as exploratory, and we say so here instead of rounding it up. Every number in the console carries the count it came from, so the class of your own programme is readable off the page.
Sample size disclosure
meets decision-gradeThe standard: Decision-grade requires the provider to disclose responses per query.
Every rate renders through one gate: a dash at zero runs, raw counts under 8 runs, and a rate with a Wilson 95 percent interval at 8 or more. Per-prompt answer counts are visible in the console and in exports.
Cadence
meets decision-gradeThe standard: Monthly or quarterly cadence is directional. Weekly or more frequent is decision-grade.
Every tracked prompt rides the weekly sweep, and the plans that fund it add a daily tripwire on the contested quarter of the pool plus a burst window when a verdict flips. The fast lane runs on a per engine clock of 1 to 5 days set by the measured survival of each engine's answers, from consumer ChatGPT daily out to Claude every fifth day. One correction we owe this table: no marketed tier includes a priority slot, so what a buyer on the current ladder actually gets is the weekly sweep, the tripwire and the burst. That is still weekly or better, which is what the criterion asks.
Reproducibility
partial, statedThe standard: Decision-grade defines acceptable variation ranges within a 7-day window, reports confidence levels, and maintains a documented process.
Confidence intervals ship on every rate at n of 8 or more, and we publish per engine stability between two runs of the same prompt from our own production data. The constants the product applies were measured on 27 August 2026: mean source overlap runs from 26.9 percent on Google AI Mode to 82.4 percent on Perplexity. The recomputation we published a week later on a larger corpus moved those to 23.9 and 88.7, which is both why one threshold across engines would mislabel most of them and why every figure here carries the day it was measured. The acceptable variation band is the part we do not have: nothing in the product can mark a prompt as a control outside the demo workspace, and no code computes a band from one, so there is nothing accruing and nothing to render. It is on the Planned list at the foot of this page, and until it is built this row is a partial.
Multi-platform reporting
meets decision-gradeThe standard: Results must be reported per platform, with any weighting disclosed.
Every rate is per engine. The one pooled headline covers only consumer buyer surfaces, names the developer API surfaces it excludes and how many answers that was, and each engine is reported on its own beside it. Every engine is labeled by the surface actually measured (consumer scrape or API), because the two disagree: asked the same commercial questions inside the same minute, the ChatGPT API and chatgpt.com shared one domain of seven on the first question, and on the second shared only Trustpilot and Reddit with no recommended service in common. That study is three questions on one day, which is enough to show the surfaces can diverge and not enough for a rate, so we publish no rate for it. Claude has no measurable consumer surface at all, so any Claude number, ours or anyone's, is an API number; ours is labeled as one.
What we refuse to build, and why
All 8 refusals below are enforced in code and tests, not just promised in copy. Competitors ship most of these; we think each one sells a number that is not real, and a tool you trust with strategy cannot do that even once.
- A rank number for AI answers. An AI answer is a sample from a distribution, not a ranked list: the same prompt asked twice returns different names in a different order (a measured test put the overlap between two identical scans an hour apart at 3.6 percent). We report share of appearance with a sample size and a range instead. A repo-wide test fails the build if rank language for AI answers ever reaches a customer surface.
- Revenue per search query. Search Console has no revenue and Analytics has no query, so any per-query dollar figure is a model presented as a measurement. We show revenue per page (Analytics' own number) with the queries that feed the page as demand context, clicks only.
- A backlink index. Editorial mentions correlate with AI citations far more strongly than backlinks do. Link data appears in exactly two supporting roles: weighing the authority of citation targets, and finding pages that link to the sources engines already cite. Selling link counts as AI visibility would be selling the weaker signal.
- "AI answers" served from model APIs alone. Our own divergence study found the consumer ChatGPT surface recommending almost entirely different providers than the API surface on the same questions, inside the same minute. It is three questions on one day and it is published with that limit written on it, which is why it reports no rate. We measure both surfaces, labeled separately.
- One visibility number averaged across developer APIs and consumer apps. The two shared almost no sources on the questions we compared, so an average across them describes an answer nobody receives, and a brand could appear to improve on a surface none of its buyers read. Our pooled line covers only the surfaces a buyer can actually reach, it names the developer APIs it left out and how many answers that was, and each one is reported on its own beside it.
- Schema as a citation lever. Structured data is hygiene: it helps machines parse your facts without guessing. No measured evidence shows it makes AI engines cite you, so our free schema generator says exactly that on the tin.
- AI-estimated search volume for prompts. Vendor "AI search volume" figures are derived from People Also Ask and autocomplete data, not from observed AI usage. Where we show demand, it is real search volume labeled as search volume, or our own measured answer rates.
- Numbers without denominators. Every rate in the product carries the sample it came from. A failed engine call is a failed run, excluded from the denominator, never counted as an answer that ignored you. An unmeasured value renders as a dash, never a zero.
How we grade ourselves
Every tool in this category says its numbers are accurate. These are ours, scored against hand labeled sets of real AI answers, one per stage of the pipeline. Every stage is graded on demand, by a person running the eval command: no build step runs them, the deterministic ones included, and the model backed ones cost money to grade. That is why the date of the last run is printed below, and why it goes stale rather than refreshing itself.
Last graded 2026-09-06, across 25 labeled cases. A perfect score on a small set is a small claim, which is why the case count sits beside every number. The sets grow from real production answers, so they get harder over time, not easier.
Planned, not built yet
None of the 5 items below is built. They are the places where the honest version of a rule on this page is still ahead of the code, so they are listed here instead of being described upstairs as though they shipped. Each one says what happens instead today.
- Intervals that account for clustering by prompt, day and engine. Today every pooled interval treats each answer as an independent draw, on a corpus where we measured design effects of 2.8 to 4.4.
- Reallocating runs between cells on the weekly cadence. Today the planner reallocates nothing at one run per prompt per engine, which is what every scheduled sweep asks for. It still caps what a cycle may spend.
- Control prompts you can mark, and the acceptable variation band measured from them. Today you cannot mark a prompt as a control outside the demo workspace, and no code computes a band, so the IAB reproducibility row stays a partial.
- Holdout readouts that settle on their fixed days and then stop moving, the way the weekly client verdict already does. Today a running experiment's verdict is recomputed on every read over a window ending today, so it can read positive on day 12 and turn by day 30. Only a concluded experiment is settled.
- Held out questions the product actually protects, and receipts judged on the named rate as well as the cited rate. Today nothing blocks a brief, a draft or a placement on a held out question, and a receipt's baseline and post are both the cited rate.
Known limits
- · Retrieval-query (fanout) capture covers only engines that expose their queries in API metadata today. Absence means not exposed, never no retrieval.
- · Attribute extraction is model-assisted; we store a row only when its evidence sentence exists verbatim in the answer and names the subject. The receipt is shown on hover.
- · Mention order within an answer is reported as order of naming, not as a rank claim.
- · Sample sizes are what they are: small run counts produce wide intervals, and we show the interval instead of hiding it.
The standing challenge: run Ansengine and any other tool on the same brand for the same week, and compare what each is willing to put a confidence interval on.
