A defensible AI brand sentiment score
Replace easy-to-game AI brand sentiment totals with attribute coding, blind review, repeated samples, reviewer agreement, and uncertainty.
A defensible AI brand sentiment score starts with preserved answers, predefined attributes, blind independent coding, repeated cross-product samples, mixed and uncertain labels, and reviewer agreement. Report the components and denominators so readers can see what the summary represents.
A sentiment total is easy to manipulate. Add more branded praise prompts, remove mixed answers, count every mention as positive, or rerun until the output improves. A polished decimal does not repair a biased sample.
The limits of one sentiment score
Generated answers can praise one attribute and criticize another. A product may be framed as capable but expensive, secure but complex, or suitable for enterprises but not small teams. Collapsing that into one polarity discards the buyer-relevant tradeoff.
Sentiment belongs beside other measures rather than inside a single blended score. Confirm the entity, then code evaluative attributes. Keep recommendation, citations, and observed business outcomes in separate report columns. The AI brand perception audit defines the full measurement ladder and its denominators.
General sentiment analysis references explain why task definitions matter. The Stanford Sentiment Treebank research models fine-grained, compositional sentiment rather than treating all documents as simple polarity. NIST’s AI Risk Management Framework 1.0 describes validity, reliability, transparency, and documented limitations. These principles support a careful method. They do not set a benchmark for a brand score.
Build an attribute taxonomy before collection
Choose attributes from buyer decisions and company risk. A B2B software taxonomy might include:
- capability and performance
- ease of implementation
- interoperability
- security and compliance
- reliability
- pricing and value
- service and support
- customer and persona fit
- limitations and risks
- evidence quality
Define each attribute and add inclusion and exclusion examples. Freeze the taxonomy before reviewers see the main sample. Adding a favorable attribute afterward creates researcher discretion.
Use a fixed prompt panel with neutral category, comparison, objection, branded fact, and persona-specific groups. The ChatGPT brand sentiment analysis guide explains how to distinguish observing generated sentiment from using ChatGPT to label customer text.
Create a blind coding rubric
For every evaluative clause, record entity, attribute, exact passage, polarity, factual accuracy, materiality, citation, and confidence. Use these labels:
- Positive: clearly favorable framing for the stated persona.
- Negative: clearly unfavorable framing or risk.
- Neutral: descriptive without meaningful evaluation.
- Mixed: linked favorable and unfavorable framing that should not be separated.
- Uncertain: insufficient context, ambiguous entity, or unclear evaluation.
Do not use neutral as a bin for disagreement. Do not force mixed claims into the stronger adjective. Preserve caveats because they often contain the most useful buyer information.
Blind reviewers to company preference where practical. Randomize answer order and remove dashboard labels that reveal the desired result, while retaining enough context to judge persona fit. Never remove evidence that changes meaning.
Run repeated cross-product samples
Record product, visible mode, market, account state, exact prompt, date, and run number. Run identical prompts across products only when the modes are reasonably comparable, and report each product separately before any combined view.
Repeat high-value prompts according to a policy set in advance. Save refusals, absent brands, and uncited claims. Excluding them can bias the report. Do not compare a browsing answer with an offline answer as though only the model name changed.
The sampled answer is the unit of observation. It is not a census of what all users see. Prompt wording, account context, product updates, and stochastic generation can change results.
Measure agreement and uncertainty
Have at least two reviewers independently code a calibration subset. Discuss disagreements, refine definitions, then re-code the subset before the main sample. Keep the original disagreement record.
Report:
- raw agreement by attribute and label
- a confusion table showing which labels reviewers mix up
- unresolved cases and adjudication rules
- sample size and missing answers
- variability across repeated runs
- reviewer confidence distribution
Chance-corrected measures such as Cohen’s kappa can be useful for two reviewers, although prevalence and small samples can make one statistic misleading. Publish raw agreement and the underlying counts. The original Cohen kappa paper record defines the chance-agreement concept. Each study still needs a justified decision threshold.
Confidence intervals can describe sampling uncertainty when the sampling design and assumptions justify them. They do not capture prompt bias, coding errors, product changes, or missing parts of the prompt universe. State those limitations separately.
Use a sentiment scoring workbook
Create these workbook tabs:
- Panel: prompt IDs, groups, personas, weights fixed in advance.
- Runs: conditions, complete responses, citations, missing status.
- Claims: one evaluative clause per row with attribute and evidence.
- Reviewer A and B: independent labels and confidence.
- Agreement: raw tables, optional kappa, disagreements, adjudication.
- Volatility: attribute and polarity changes across repeated runs.
- Report: raw counts, rates, limitations, and material findings.
If stakeholders insist on a summary rate, publish the formula, denominator, weights, and components. Show “18 positive, 7 negative, 6 mixed, and 4 uncertain attribute claims across 35 coded claims” alongside the net total. Rank competitors only after applying identical prompts and coding.
Report components instead of vanity totals
Break results down by attribute, persona, prompt group, product, and period. Pair every material conclusion with passages and sources. Keep factual accuracy separate: a favorable claim can be false, and an unfavorable caveat can be accurate.
Avoid causal language. If sentiment improves after a source correction, say it improved in the sampled runs after the correction. Retrieval and generation may have changed for other reasons. Avoid ranking guarantees and universal “good score” thresholds.
For issues that are materially false rather than merely unfavorable, follow Leaf’s wrong company information protocol. To audit the site, evidence, and measurement system together, start with Leaf’s AEO assessment, then track fixes through repeated samples.
Frequently asked questions
Can AI be used for brand sentiment analysis?
Yes, as a classifier or as the object being measured, but those are different studies. Automated labels should be validated against a defined human-coded task. Generated brand sentiment requires controlled prompts and preserved answers.
What is a good sentiment analysis score?
A good score is reliable enough for its decision, transparent about the sample, accurate at attribute level, and accompanied by raw counts and uncertainty. Compare only consistent methods.
Why is positive, neutral, and negative scoring insufficient?
It hides mixed attributes, persona-dependent fit, uncertain entity identity, factual errors, and caveats. Diagnose these losses by coding exact clauses and comparing reviewer disagreements instead of forcing one document-level label.
How should mixed statements be scored?
Keep mixed as its own label when favorable and unfavorable framing are meaningfully linked. Record the attribute and exact passage. Split only when clauses evaluate distinct attributes and can be interpreted independently.
How do you measure agreement between sentiment reviewers?
Have reviewers code the same preserved sample independently, then report raw agreement and a confusion table. Add a chance-corrected statistic when appropriate, with counts and assumptions. Investigate disagreements before trusting aggregate results.
Should AI brand sentiment reports include uncertainty or confidence intervals?
Yes, when they are clearly defined. Include run variability, reviewer confidence, unresolved labels, and sampling intervals where assumptions support them. A confidence interval does not account for a biased prompt panel or future product changes.