Measure LLM recommendation quality and share of voice
Measure LLM recommendation quality alongside share of voice using transparent prompts, rationale, fit, caveats, factuality, and citation support.
LLM share of voice measures how often a brand appears within a declared prompt sample. Pair it with a disclosed recommendation-quality scorecard to show whether each mention is accurate, suitable for the buyer, supported by evidence, and positively framed. Retain the underlying answers so reviewers can reconstruct the result.
Define the denominator and buyer context
Calculate mention share as brand-mention runs divided by eligible runs in a frozen panel. Publish both numbers. Separate branded factual prompts from neutral category and comparison prompts, because asking directly about a company makes inclusion easier. Group prompts by awareness, evaluation, objection, and implementation stage, with persona and market recorded.
A quality scorecard should keep components visible rather than hide them in one weighted number:
| Component | Review question |
|---|---|
| Inclusion and order | Was the brand included, and where in an explicitly ranked list? |
| Buyer fit | Does the stated use case match the prompt’s constraints? |
| Rationale | Are reasons specific, relevant, and verifiable? |
| Caveats | Are material limitations represented fairly? |
| Factuality | Do buyer-relevant claims match dated sources? |
| Citation support | Does each displayed source entail the associated claim? |
| Framing | Is sentiment positive, neutral, mixed, or negative by attribute? |
Do not label prose order as rank when the answer is not a ranked list. Code recommendation as explicit, conditional, mentioned-only, excluded, or warned-against. Then annotate the reason. Two reviewers should resolve ambiguous cases using a written rubric rather than retrofitting labels to improve a score.
NIST’s AI Risk Management Framework 1.0 emphasizes documented measurement and context, useful principles for avoiding opaque evaluation. Google’s AI features documentation also describes query fan-out, one reason a single prompt may involve varied supporting searches. A share-of-voice benchmark still has to be defined for the specific panel and decision.
Connect quality to buyer decisions
High mention share with weak fit suggests positioning or source clarity work. Strong recommendations with unsupported claims create reputation risk. Competitor citations for a capability you document well may justify a source-entailment review, not a promise that rewriting your page will win citations. Compare website citation tracking with the cross-model brand test. Leaf’s AI search visibility measurement guide provides denominator rules.
Business value remains downstream: qualified visits, assisted opportunities, and buyer feedback. Do not multiply sampled mentions by an invented market volume. For a scoped baseline and remediation priorities, start the AEO assessment.
Add a recommendation-quality record to every eligible run
A mention should open a review, not finish one. For each eligible response, preserve the sentence or list item containing the brand and code its role. Was it an explicit recommendation, a conditional option, a passing example, an exclusion, or a warning? Then capture the buyer constraint the answer used: budget, geography, implementation capacity, integration, compliance need, or company size. This makes it possible to distinguish broad visibility from useful fit.
Evaluate rationale at the claim level. “Good for enterprises” is too vague unless the response identifies a relevant capability that can be checked. Give full credit when the rationale is specific to the prompt and supported by dated evidence. Mark generic praise separately from false claims. A flattering fabrication is a quality failure even if it raises sentiment.
Caveat scoring should ask whether the answer includes limitations that could change the decision. A response can omit a material deployment restriction while adding harmless boilerplate such as “compare your options.” A relevant limitation also need not make the framing negative. Code completeness and valence as separate fields.
Use a scorecard without hiding disagreement
Create one row per product, prompt, and repetition. Retain columns for mention, position when explicitly ranked, recommendation status, buyer fit, rationale specificity, caveat completeness, factuality, citation entailment, and attribute-level framing. Add the exact excerpt and reviewer confidence beside each judgment. If a team wants a composite score, publish the weights and the component values so the result can be reconstructed.
Calibrate reviewers before scaling collection. Give two reviewers the same small set, compare labels, and discuss cases such as an unordered shortlist, a recommendation that violates one constraint, or a citation that supports only half a sentence. Update the rubric with examples. Resolve disagreement with an adjudicated result while retaining the original reviewer notes.
Report raw counts before percentages. “Recommended in 6 of 20 eligible runs” is more interpretable than a rounded index, especially when prompt groups are small. Keep ineligible failures in an exclusions table with reasons. If one product fails to answer several difficult prompts, silently removing those runs can make its quality look better.
Read the patterns as different business problems
High mention share plus low buyer fit often indicates category recognition without precise positioning. Review the language used on comparison, use-case, integration, and qualification pages. Low mention share plus strong quality when present may instead point to discoverability or source breadth. Frequent recommendations paired with unsupported claims demand factual remediation before promotion. Amplifying them would increase reputation exposure.
Break results down by buyer stage and persona. Discovery prompts test category association, while evaluation prompts expose capability and constraint handling. An implementation answer can be excellent without naming every category leader. Do not compare those denominators as if they represent the same opportunity. The scorecard should reveal where the brand helps a specific buyer make a decision.
Competitor comparison also needs symmetry. Apply identical entity rules, prompt exposure, factual standards, and reviewer thresholds to each named brand. A competitor should not lose points because its caveat is known internally but absent from the cited public record. The evaluation concerns what the answer says and what public evidence supports, not the reviewer’s private knowledge.
Govern the metric around decisions
Assign remediation by failed component. Entity and factual errors go to the owners of canonical information. Weak rationale may call for clearer proof and use-case content. Missing material caveats require product and compliance input. Citation mismatch belongs in a source-entailment backlog. Prompt design and label disputes belong to the measurement owner, not the content team.
Trend only panels that remain comparable. Version prompt wording, buyer mix, model or product labels, collection method, and rubric. When one changes, annotate a break rather than presenting a smooth historical line. Repeated output varies, so a movement in a small sample should trigger inspection of underlying rows before an executive claim.
Connect this work to business evidence cautiously. Referral visits, sales-call mentions, and assisted opportunities can be monitored alongside the sampled answers, but the scorecard does not establish attribution. Its immediate value is sharper: it shows whether visibility is accurate, relevant, and responsibly supported for the buyers represented in the panel.
Aim for a defensible distribution of accurate recommendations for appropriate buyers, with visible limitations and reasons that survive source review. Preserve low scores and contradictory runs because they make later improvement claims credible.
Frequently asked questions
What does AI share of voice mean?
It is the proportion of eligible sampled responses in which a brand appears, relative to the declared denominator and comparison set.
How is AI share of voice calculated?
Lock a prompt panel, run each eligible test under recorded conditions, and divide brand-mention runs by all eligible runs. Publish raw counts and exclusions.
What is a good AI share of voice for a brand?
The useful percentage is the one measured on a stable panel against relevant competitors, buyer stages, and your own baseline. Inspect recommendation quality alongside it.
Why is mention share not the same as recommendation quality?
A mention may be incidental, negative, inaccurate, or a warning. Recommendation quality evaluates fit, rationale, caveats, facts, and support.
How should recommendation rationale and caveats be scored?
Use a written rubric, preserve quoted evidence, rate specificity and material completeness separately, and have reviewers resolve ambiguous cases.
Can high share of voice coexist with negative brand framing?
Yes. A frequently mentioned brand can be described critically or described as unsuitable, which is why sentiment and recommendation must remain separate metrics.
Optimize for useful recommendations
Re-run the same scorecard after material source or positioning work and inspect component-level changes. A credible program values accurate buyer fit over the largest possible mention percentage.