How to audit AI brand perception
Run a repeatable AI brand perception audit across recognition, framing, sentiment, recommendations, citations, accuracy, and business value.
In short: An AI brand perception audit is a controlled review of what named AI products say about a company across a fixed prompt panel. It records recognition, factual accuracy, framing, sentiment, recommendations, and citations separately. The results describe the sampled answers. They cannot reveal a model’s private reasoning or predict future rankings.
The phrase “AI brand perception” is often used for two different subjects: how consumers feel about a brand that uses AI, and how an AI-generated answer represents a brand. This audit addresses the second. Its unit of evidence is an observed answer under recorded conditions.
Separate the dimensions before scoring
A useful audit keeps eight questions apart:
- Access: Can a relevant page be fetched and rendered under the controls you set?
- Retrieval: Did the answer product appear to use a source for this prompt?
- Entity understanding: Did it distinguish the company, product, parent, category, and namesakes?
- Mention: Did the answer name the brand?
- Citation: Did it link to a source, and does that source support the nearby claim?
- Sentiment: Was a particular attribute framed positively, negatively, neutrally, or with mixed evidence?
- Recommendation: Was the brand proposed as suitable for the stated buyer and constraints?
- Business value: Did any observed referral or conversion follow, without assuming the answer caused it?
Treat these as distinct observations rather than a guaranteed funnel. A crawler request records access. A critical mention still counts as a mention, and a citation may support a definition while the recommendation rests on other material. Leaf’s guide to an AI visibility audit explains why each field needs its own denominator.
Official platform guidance helps define the observable boundaries. OpenAI publishes controls for its web crawlers in its crawler documentation, while Google explains eligibility for its AI search experiences in Search Central guidance. These documents explain access and eligibility, without offering a formula for predicting inclusion or citations.
Build a controlled 30-prompt panel
Create six groups of five prompts. Use actual buyer language and preserve exact wording:
- identity and canonical company facts
- category and problem recognition
- audience and use-case fit
- comparison and alternatives
- objections, risks, and limitations
- proof, implementation, and purchasing criteria
Keep neutral discovery prompts separate from branded checks. “What does Acme do?” tests recognition and accuracy. Whether Acme appears when a buyer asks for category options is a separate result. Include persona and constraint details only where they reflect a real buying situation. Keep recommendations for a small local firm separate from those for a regulated global enterprise.
When the question is specifically whether buyer context changes the recommendation, use the persona-dependent AI recommendation test rather than duplicating its one-variable-at-a-time matrix here.
Before collection, record the product, visible mode, signed-in state, market, date, panel version, and repetition policy. Logged-in experiences may use account context or product features that anonymous sessions do not. Therefore, compare like with like rather than treating one session as universal.
For high-value prompts, repeat runs on a declared schedule. Variability is itself a finding. Do not quietly rerun until the preferred answer appears. Preserve absences, refusals, and incomplete answers alongside favorable results.
Use an LLM Brand Perception Card
Create one row per prompt run and include these columns:
| Field | Record |
|---|---|
| Context | Product, mode, date, market, account state, run number |
| Prompt | Stable ID, exact text, intent, persona |
| Entity | Correct entity, collision, ambiguity, or no recognition |
| Presence | Absent, mentioned, compared, or recommended |
| Framing | Attribute, polarity, caveat, and exact supporting passage |
| Accuracy | Correct, incorrect, ambiguous, stale, or unverifiable |
| Sources | Every cited URL and the claim each appears to support |
| Value | Buyer relevance and separately observed site outcome |
| Confidence | Reviewer confidence and reason |
The card is the audit’s central artifact. If executives need a summary, show raw counts beside rates: “recommended in 4 of 30 recorded runs.” A label such as “13% of AI” claims a scope the sample cannot support. Publish the denominator, prompt mix, and missing runs.
For sentiment work, code claims at attribute level. An answer can praise ease of use while warning about enterprise controls. Calling that response simply “positive” destroys useful information. The ChatGPT brand sentiment method provides a dedicated annotation approach, while the defensible sentiment scoring guide covers reviewer agreement.
Compare models and repetitions fairly
Do not compare a browsing mode in one product with a non-browsing mode in another and attribute the difference to model quality. Report each tested experience as configured. Save complete answers and citations because interfaces and source displays change.
Review disagreement at three levels: between repeated runs in one product, between products, and between reviewers. A stable factual error deserves different action from one isolated omission. A citation that appears repeatedly may indicate an important source path, but it does not expose internal weighting.
Competitors should use the same prompts and coding rules. This controls the most obvious source of bias: testing your brand with generous branded prompts while testing competitors only in neutral discovery. The competitor recommendation audit follows the full shortlist from inclusion through rationale and caveats.
For a complete product-to-product protocol, use the ChatGPT, Perplexity, and Gemini brand test and keep product modes and conditions explicit.
Turn findings into a bounded action plan
Prioritize issues that could change a buyer’s decision. Wrong compliance claims, discontinued products, and confused namesakes belong near the top. A missing slogan usually belongs in routine monitoring.
Classify each action:
- repair robots, status, canonical, rendering, or discovery problems you can verify
- make canonical company facts direct, dated, and internally consistent
- correct stale profiles or independent sources through their normal editorial process
- add primary evidence where buyers need proof
- submit platform feedback when an interface offers it
- monitor uncertain or one-off observations before changing content
Retest the original panel after changes. A later improvement is an observed sequence, not proof that one edit caused it. Retrieval, model behavior, source availability, and presentation can all change.
For a structured starting point, run Leaf’s AEO assessment. Use the results to scope access, content, entity, and measurement checks, then verify the proposed work against recorded runs.
Frequently asked questions
What is an AI brand perception audit?
It is a documented sample of how named AI products represent a brand across fixed prompts and conditions. It separates factual representation, sentiment, recommendation, and sources instead of treating a mention as success. Results apply to the sampled runs and declared conditions.
Why should I track AI brand visibility?
Track it to find material inaccuracies, weak category recognition, competitor source gaps, and buyer questions your evidence leaves unanswered. Diagnose access, retrieval evidence, entity accuracy, mention, citation, and recommendation separately. The report should stay within observable evidence.
Who inside the organization should own AI brand perception?
A cross-functional owner in SEO, communications, or product marketing should coordinate it, with web, analytics, legal, and subject experts handling issues they control. One accountable lead prevents conflicting corrections. Ownership structure depends on company risk and resources.
Does using a logged-in account change the answers captured?
It can. Account context, memory, location, settings, and available modes may affect the observed experience. Record account state and compare consistent conditions. A logged-in result should not be generalized to anonymous users.
Should competitors be audited at the same time?
Yes, when the decision concerns category visibility or recommendation quality. Apply identical neutral prompts, repetitions, and coding. Competitor comparison provides context, but a sampled advantage is not a permanent market ranking.
How often should an AI brand perception audit be repeated?
Establish a baseline, monitor high-risk prompts monthly or quarterly, and retest after material company, source, or access changes. Increase frequency for volatile answers or safety-critical facts. Cadence should follow risk and observed variability, not a universal schedule.