AI search visibility metrics and KPIs
The metrics worth reporting — mention rate, share of voice, citation share, average position, sentiment — and the ones that mislead.
Reporting on AI search visibility is harder than reporting on rankings, because there is no established convention and every tool names things differently. This article sets out the metrics that carry real information, what each one is good for, and which ones create false confidence.
The core metrics
Mention rate. The share of tracked responses that name your brand, per engine. This is the base metric and the one most dashboards lead with. It is only interpretable alongside the prompt set it was computed from, so always report the two together.
Share of voice. Your mentions as a proportion of all brand mentions across the same responses. This is more robust than mention rate because it normalises for prompt difficulty — if every brand's mention rate falls because the engine got terser, share of voice stays stable and correctly shows no change in your relative position.
Citation share. How often your domain is used as a source, as a proportion of all cited domains. This measures something different from mention rate, and the gap between the two is diagnostic. High mentions with low citations means the model knows you but is not reading your site. Low mentions with high citations means your content is useful but your brand is not being credited.
Average position. Where in the response your brand first appears, averaged across responses. First-named brands carry disproportionate weight with readers, so a rising mention rate with a falling position is a weaker result than it looks.
Sentiment. The distribution of favourable, neutral, and cautionary mentions. Report it as a distribution, not an average, because an average hides a bimodal pattern where half the engines recommend you and half warn about you.
Competitor gap. The difference between your mention rate and each tracked competitor's on the identical prompt set. In practice this is the single most useful number for a leadership report, because it is stable, comparative, and immune to the engine-wide shifts that move everyone's absolute score at once.
Supporting metrics
- Prompt coverage — the share of your tracked prompts where you appear at least once. Distinguishes broad weak presence from narrow strong presence.
- Engine spread — how evenly your visibility is distributed across engines. Concentration in one engine is a risk.
- Cited page mix — which of your URLs are being used. Tells you what to produce more of.
- AI crawler activity — how often AI crawlers fetch your site. A leading indicator: crawl growth usually precedes citation growth.
Metrics that mislead
A blended cross-engine visibility score. Averaging across engines that behave completely differently produces a number that moves for reasons you cannot diagnose. Keep engines separate, and blend only for an executive summary.
Any single-run result. One favourable answer is an anecdote. Screenshots are useful for illustration and useless as evidence.
Week-over-week deltas on small prompt sets. Below roughly thirty prompts, the noise exceeds most real movements, and you will spend meetings explaining random walks.
Raw mention counts. Counts move with how many times you ran the prompts. Always report rates.
An honest reporting pack
- Mention rate and competitor gap per engine, trended over at least eight weeks.
- Share of voice against a fixed competitor set.
- Citation share plus the top cited pages, yours and theirs.
- Sentiment distribution, with examples of the worst framings.
- A note of any prompt set changes in the period, because these invalidate comparisons.
- One or two verbatim answers, included as illustration and labelled as such.
Common questions about AI visibility KPIs
- Which single metric should I report to leadership? Competitor gap on a frozen prompt set. It is stable, comparative, and immune to the engine-wide changes that move everyone's absolute score at once.
- What is a good AI visibility percentage? There is no universal benchmark, and anyone offering one is guessing. Around 30% on narrow high-intent B2B prompts is strong; the same figure on broad consumer questions would be exceptional. Judge against your competitor set, not an industry number.
- How do I set a target? Per question cluster rather than overall. Definitional questions are easy to win and rarely drive revenue; comparison questions are hard to win and frequently decide deals. One blended target hides that entirely.
- Should I include sentiment in the KPI set? Yes, as a distribution rather than an average. An average conceals the case where half the engines recommend you and half warn about you.
- How far back should the trend go? At least eight weeks before drawing conclusions. Below six weeks you are mostly reading variance.
- What do I do when the number drops? Open the stored responses and read what changed before doing anything else. Most drops trace to a prompt set edit, a model update, or a competitor publishing something better — and the response differs for each.
Tip: agree the prompt set with whoever receives the report before the first one is sent. Every argument about AI visibility numbers is eventually an argument about which questions were counted.
Still stuck? We typically reply within 1 business day.
Contact support