Insights
    AI Visibility
    Measurement
    GEO

    How to Measure Your Brand's AI Visibility

    Avisible TeamJune 16, 20266 min read

    Most companies have analytics for website traffic, search rankings, ad performance and conversions. Very few can say how they appear inside AI systems. That is becoming a significant blind spot: as ChatGPT, Gemini, Claude and Perplexity shape more of how customers discover brands, the channel with no dashboard is the one quietly deciding who gets shortlisted.

    The good news is that AI visibility is measurable. Not perfectly, and not with the precision of a rank tracker — but well enough to make decisions with, and far better than the spot-checking most teams currently do.

    What AI visibility actually means

    AI visibility is whether AI systems mention your brand when someone asks a question you should own, how often competitors appear instead, what those systems say about you, and which sources they rely on to say it.

    It is not a single score. Treating it as one number is the fastest way to produce a report nobody can act on, because the interesting information is always in the breakdown: which stage of the buying journey you disappear at, which competitor takes your place, and which source is feeding the description you dislike.

    Why your SEO metrics will not tell you

    The instinct is to assume strong search performance implies strong AI performance. The data says otherwise. Ahrefs found only around 12% of URLs cited by AI assistants also rank in Google's top ten — meaning most of what AI cites is not what search ranks first.

    The reason is that the two systems reward different things. Search rewards relevance and link authority for a page. AI systems weight entity clarity, corroboration across independent sources, and how directly a passage answers a question. A page can be excellent by one standard and unusable by the other. If your brand ranks well and still goes unmentioned, the reasons are usually structural rather than mysterious.

    The metrics worth tracking

    Five carry most of the signal:

    1. Mention rate. How often your brand is named across a defined set of buyer questions. The base number everything else is read against.
    2. Share of answer. How much of the response is about you compared with competitors. Being mentioned last in a list of six is not the same as being the recommendation.
    3. Sentiment and accuracy. What AI says about you, and whether it is true. Confidently stated wrong pricing does more damage than absence.
    4. Entity clarity. Whether the system knows who you are without confusing you with a similarly named company. This one predicts the others.
    5. Citation sources. Which pages AI leans on when describing your category. Usually the most immediately actionable output, because it tells you where the work has to happen — often on sites you do not control.

    The method matters more than the metrics

    Any of those numbers can be produced badly. Three things separate a measurement you can act on from a screenshot of one chat session.

    Ask real buyer questions, not brand questions

    Asking an AI "what is [your brand]?" tells you almost nothing commercially useful — you have named yourself in the prompt, so of course you appear. The questions that matter are the ones buyers ask before they know you exist: category questions, comparison questions, problem-framed questions. Designing that set is most of the work, and we go through the types in the questions every brand should test in ChatGPT.

    Run every question more than once

    AI answers are not deterministic. The same question asked three times can return three different sets of company names. A single run is an anecdote; repeated runs turn it into a rate you can compare over time. Our own Snapshot runs each of 100 questions three times across three systems, which is where the figure of 900 answers comes from — the repetition is the point, not the volume.

    Test more than one system

    ChatGPT, Gemini, Claude, Copilot and Perplexity do not agree with each other. They ground answers in different sources and weight authority differently, so a brand can be well covered in one and absent from another. Measuring a single system produces a confident conclusion about a fraction of the market.

    The mistakes that waste the effort

    Four recur often enough to be worth naming:

    1. Prompting from your own account. Personalization and memory can surface your brand to you and nobody else. Test logged out.
    2. Measuring in English only. If your buyers ask in Finnish, Swedish or German, an English-only measurement misses the market you actually sell in — and AI systems behave differently across languages, as the English-bias pattern shows.
    3. Counting mentions by naive string match. In inflected languages a brand name changes form inside a sentence, so simple substring matching undercounts real mentions. This is a genuine measurement bug, not a rounding error.
    4. Measuring once. Models update. A baseline with no second reading cannot tell you whether anything you changed worked.

    From measurement to decisions

    A measurement is only worth the action it enables. In practice the output should answer three questions: where are we strong, where are we exposed, and what is the highest-value thing to change first.

    That last one usually turns on citation sources rather than your own content. If AI describes your category using three industry roundups you do not appear in, publishing another page on your own site will not move it. The lever is the source, not the page — a dynamic we explored in how AI agents decide which brands to recommend.

    How often to re-measure

    More often than annually, less often than weekly. The useful cadence is quarterly for most companies, with two exceptions.

    Measure again sooner if you have shipped something material — a repositioning, a new market, a name change, a significant piece of coverage. Those are the moments when the number should move, and checking tells you whether the work reached the systems that matter.

    Measure again sooner if a major model updates. Answers can shift noticeably when a system changes how it grounds responses, and a shift you did not cause is still a shift you need to know about.

    What does not justify a re-run is a bad week. Because answers vary between runs, short-interval monitoring mostly measures noise, and teams that watch a daily dashboard tend to react to variance rather than to trend. If you are going to track continuously, track the rate across a fixed question set rather than individual answers.

    Getting a baseline

    You can run a version of this yourself: write thirty real buyer questions, ask each three times in two systems while logged out, and record which companies get named. It is tedious and it is not rigorous, but it will tell you more than you know today.

    If you would rather have it done properly and comparably, that is what our Visibility Snapshot is — a fixed set of buyer questions, run repeatedly across three systems, counted by machine and read by a consultant. Either way, the important step is the first one: find out what AI already says about you, because it is saying something whether you have looked or not. Common questions about scope and cost are answered on our FAQ page.