Measurement
How do you measure your brand's visibility in AI answers?
Last reviewed: · By Kraffic.ai
Short answer
You measure AI visibility by running a fixed set of buyer prompts on each AI assistant, saving the answers, and counting how often your brand is named. The core measures are a visibility score (the share of answers that name you), share of voice against competitors, per-assistant results, and prompts won or lost since the previous scan. Repeat it on a schedule, because answers vary.
Key takeaways
- AI visibility is measured by sampling: run prompts, save answers, count mentions.
- A visibility score is the share of answers that name your brand, so its meaning depends on which prompts are tracked.
- Share of voice compares your mentions with those of competitors you have approved.
- Report each assistant separately, because results on ChatGPT, Claude and Gemini often differ.
- Because answers vary between runs, trends over repeated scans are more reliable than any single result.
- A number is only useful if you can open the answer it was calculated from.
Prompts this guide answers
- “how to measure AI visibility”
- “AI visibility score”
- “track brand mentions in ChatGPT”
- “AI share of voice”
- “How can I find out how often ChatGPT, Claude and Gemini mention my company compared with my competitors?”
- “What metrics should I report to my CEO to show whether our brand is visible in AI search?”
- “AI search visibility tracking tool”
- “Why do I get a different answer every time I ask an AI assistant about my industry?”
What is AI visibility?
AI visibility is how often and how prominently AI assistants name your brand when people ask questions relevant to your business. It is the GEO counterpart to search visibility.
It cannot be read from a report supplied by the AI companies, because none of them publishes which prompts mention which brands. It has to be measured from outside by asking the assistants and recording what they say.
Which metrics matter?
A small set of metrics covers most decisions. Each answers a different question, so they are best read together.
| Metric | What it measures | Question it answers |
|---|---|---|
| Visibility score | Share of AI answers that name your brand | How often do we appear at all? |
| Per-assistant results | The same measure split by ChatGPT, Claude and Gemini | Where are we strong and where are we missing? |
| Rank against competitors | Your position among approved competitors by mentions | Who is ahead of us? |
| Share of voice | Your mentions as a share of all tracked brand mentions | How much of the conversation is ours? |
| Prompts won or lost | Prompts where you newly appear or no longer appear since the last scan | What changed? |
| Citations | Websites named in the AI answers | Which sources shape these answers? |
| Sentiment | Themes and tone of what is said about you | Are we described well? |
| Factual accuracy | AI statements that contradict your confirmed facts | Are we described correctly? |
How is a visibility score calculated?
A visibility score is the number of answers that name your brand divided by the number of answers collected, expressed as a percentage. If a scan collects answers to your tracked prompts across three assistants and your brand is named in a portion of them, that portion is the score.
The score depends entirely on the prompt set. A list made mostly of prompts that include your brand name will produce a high score that says little; a list of non-branded buyer prompts is a harder and more useful test. Always ask what prompts are behind a score before comparing it with anything.
Why do AI answers vary from one run to the next?
AI answers vary because assistants generate text with a degree of randomness, so the same prompt can yield different wording and a different list of brands each time. When an assistant searches the web, the pages it retrieves can also change.
Model updates add another source of change, and they happen without notice. For measurement this means a single answer is an anecdote. Reliable conclusions come from many prompts, repeated scans and the trend between them.
How often should you measure?
Measure often enough to see a trend, and no more often than you can act on. For most businesses a regular schedule, kept consistent, is more useful than occasional bursts of checking.
Kraffic.ai runs tracked prompts on Claude, ChatGPT and Google Gemini at most once a day. Each scan is compared with the previous one, so you see prompts won or lost instead of having to compare raw answers by hand.
How do you set up measurement step by step?
Set up measurement in a fixed order so the numbers stay comparable over time.
Confirm your facts
Write down the business facts that must be right: what you sell, where, for whom, and key product details.
Approve competitors
Decide which companies count as competitors. This list defines rank and share of voice.
Build the prompt set
Collect non-branded and branded prompts, short and conversational, that reflect how buyers ask.
Run and save
Run each prompt on each assistant and keep the full answer, with the date.
Score and compare
Calculate the metrics, then compare with the previous scan.
Keep the set stable
Change prompts and competitors deliberately and note when you do, or trends become meaningless.
What makes a measurement trustworthy?
A measurement is trustworthy when you can trace it. For any score you should be able to see the prompts, the assistant, the date and the full answer, and count the mentions yourself.
In Kraffic.ai every number is traceable to the saved AI answer, sentiment claims are shown with the exact quote, and data can be exported as CSV for your own checks.
FAQ
Questions and answers
What is a good AI visibility score?
There is no universal benchmark, because the score depends on which prompts are tracked and how competitive the category is. A score is most useful compared with your own previous scans and with the competitors you approved. Be wary of anyone quoting an industry average without showing the prompts and method behind it.
What is share of voice in AI search?
Share of voice is your brand's mentions as a proportion of all mentions of tracked brands across the collected AI answers. If assistants name you and your approved competitors many times in total, your share is the part of those mentions that are yours. It shows your standing relative to competitors, not just whether you appear.
Can I measure AI visibility manually?
Yes, at small scale. Run a list of prompts on each assistant, paste the answers into a spreadsheet, and mark which brands are named. It becomes time-consuming quickly, since you need several assistants, many prompts and repeated runs. Tracking platforms automate the collection and keep the saved answers for comparison.
Do AI companies provide data on how often my brand is mentioned?
No. No AI company publishes its users' prompts or a report of which brands its assistant mentions. Any figure for prompt volume or brand mentions comes from outside measurement or estimation. That is why the method matters: ask what prompts were run, on which assistants, when, and whether the answers were saved.
Why is my brand visible on one assistant and not another?
Each assistant has its own model, training data, search system and crawlers, so each sees a different picture of the web. A brand can be well covered in the sources one assistant relies on and thin in another's. Per-assistant results, together with the websites each one cites, show where the gap is.
How do I report AI visibility to leadership?
Keep it to a few measures and their trend: visibility score, share of voice against named competitors, results per assistant, and prompts won or lost since the last report. Add any factual errors found and what was done about them. Include two or three real AI answers as examples so the numbers are concrete.