SEO AuditAEOAI SearchB2B Growth

AI visibility audit: a reproducible method

Learn how to run a defensible AI visibility audit with a fixed prompt panel, evidence capture, claim checks, competitor analysis, and practical retesting.

Leaf Team
August 4, 2026
9 min read

An AI visibility audit is a controlled snapshot of how named answer products represent your company across a defined set of buyer questions. It can show whether your brand is mentioned, which sources are cited, what competitors appear, and where factual errors recur. It cannot reveal a universal “AI ranking,” because there is no single result set shared by every model, mode, account, location, and date.

That distinction makes the audit useful rather than academic. A B2B team does not need another speculative score. It needs evidence that can be traced to a prompt, response, source URL, and collection time—then converted into work that content, web, PR, and product marketing owners can complete.

Start with the decisions the audit must support

Do not begin by opening five chat products and asking whatever comes to mind. First write down the commercial decisions the audit should inform. Typical decisions include:

This keeps the scope bounded. “See how we perform in AI” is not a workable brief. “Assess our visibility and factual accuracy across 40 security-software discovery, comparison, and objection prompts in the US during August” is.

A useful scope also names the products being tested. Google documents that AI Overviews and AI Mode use Search infrastructure and may use a “query fan-out” technique to issue related searches (Google Search Central). OpenAI separately documents the crawlers used for search and training controls (OpenAI crawler documentation). Those are verified platform facts. They do not imply that both products retrieve, rank, or cite sources in the same way.

Build a prompt panel that reflects B2B buying work

The prompt panel is the denominator for every visibility rate you report. Build it from actual buyer jobs, not from variations of your target keyword. Interview sales, review call notes, inspect site search, and examine the questions prospects ask before procurement.

Cover at least four intent groups:

  1. Discovery: “What approaches can reduce cloud security review time?”
  2. Category: “What is continuous compliance monitoring?”
  3. Comparison: “Compare continuous compliance platforms for a mid-market SaaS team.”
  4. Objection or proof: “What evidence should I require before buying a compliance automation platform?”

Keep neutral and branded prompts separate. Neutral prompts measure sampled discovery; branded prompts are better for checking whether the company description, product scope, integrations, locations, and policies are represented accurately.

Freeze exact wording for the baseline. If you improve a prompt later, create a new panel version rather than silently replacing it. Prompt wording can materially change an answer, so an undocumented rewrite breaks the comparison. For high-value questions, run multiple repetitions and record each one. Repetition does not remove variability, but it makes a one-off response less likely to dominate the conclusion.

Capture evidence at response level

A screenshot by itself is weak evidence. It is hard to aggregate, may omit citations below the fold, and often loses account or mode context. Keep screenshots when useful, but pair them with a response-level record.

Capture these fields for every run:

Field What to record Why it matters
Prompt ID Stable ID and exact prompt text Connects responses across retests
Product context Product, visible model or mode, signed-in state Prevents unlike results being combined
Collection context Date, time, market, device if relevant Establishes the snapshot boundary
Answer evidence Full answer text or complete capture Supports later accuracy review
Brand outcome Mentioned, recommended, compared, or absent Separates different types of presence
Citation outcome Linked source, domain, exact URL, claim supported Distinguishes a citation from a mention
Competitors Named companies and cited competitor URLs Reveals the comparison set
Accuracy Correct, wrong, ambiguous, or not verifiable Turns monitoring into repair work
Reviewer note Short interpretation and confidence Preserves judgment without disguising it as fact

A mention is not a citation. A citation is not necessarily a recommendation. A recommendation is not necessarily positive once caveats are read. Preserve these as separate fields instead of collapsing them into one score.

Audit access without pretending it proves inclusion

Before diagnosing content, check whether the pages you expect to support an answer are technically reachable. Review HTTP status, canonical target, meta robots, robots.txt policy, rendered content, and internal discovery paths. If your policy allows OpenAI search crawling, verify the relevant controls against OpenAI’s current bot documentation rather than copying a generic robots rule from a blog.

This work establishes eligibility conditions you control. It does not prove that an external system fetched a page, placed it in an index, retrieved it for a prompt, or used it in an answer. Server logs can provide evidence of crawler requests, but even a logged fetch is not proof of citation.

For Google’s AI search features, Google says pages must be indexed and eligible to appear in Search with a snippet; it also says no special AI text file or special schema is required. That is current official guidance, not a guarantee that an eligible page will appear. If an audit vendor proposes a proprietary markup tag as a required AI inclusion switch, ask for platform documentation.

Review citations and claims, not just brand counts

The most valuable audit work often happens after the mention count. Open every cited URL. Identify which exact part of the answer the source appears to support. Then classify the source:

When a competitor is cited, resist the reflex to imitate its page. Ask what evidence the cited page provides that yours does not: a precise definition, comparison criteria, original documentation, named methodology, implementation detail, or an independently verifiable fact. The useful output is a source gap, not “write something similar.”

Accuracy review requires the same discipline. Select material claims—pricing model, deployment method, product capability, security status, customer eligibility, geographic presence—and compare each one with the current canonical source. Mark ambiguity when scope or date is unclear. Do not label an answer false simply because its wording differs from marketing copy.

For a deeper process, use Leaf’s content audit for AI search to build a claim inventory before you retest visibility.

Use a practical triage protocol

Run this protocol after collection to turn observations into an implementation backlog:

  1. Confirm the observation. Re-run material errors and surprising absences before escalating them. Preserve both runs.
  2. Locate the expected source. Name the owned URL that should answer the prompt or verify the claim.
  3. Test the source. Check access, rendering, canonicalization, directness, evidence, date, and internal links.
  4. Check contradictions. Search product pages, documentation, press releases, profiles, and partner listings for stale variants.
  5. Classify the intervention. Choose access fix, claim correction, content improvement, third-party correction, digital PR, or monitoring only.
  6. Assign an owner and retest date. Give the work to the team that can change the underlying source.
  7. Keep confidence explicit. A confirmed technical defect may have high confidence; a proposed citation-oriented rewrite is an experiment.

A simple decision table prevents teams from prescribing content for every outcome:

Observation First investigation Likely owner
Wrong branded fact appears repeatedly Owned and third-party claim consistency Product marketing or communications
Correct page exists but is blocked or non-canonical Access and index controls Web or engineering
Competitor sources provide stronger primary evidence Evidence and source gap Content, research, or PR
Brand is mentioned but no owned page is cited Citation context and query intent SEO/content
One isolated absence Repeat before acting Analyst

Report rates with honest denominators

Report “mentioned in 9 of 40 prompts across two recorded runs” rather than “22.5% AI share of voice.” The first claim states exactly what was sampled. The second sounds like market-wide coverage even though the panel, products, and collection period are limited.

You can calculate mention coverage, citation coverage, factual accuracy, and cited-domain share within the sample. Show raw counts next to percentages. If prompts are weighted, publish the weights and business rationale. Keep downstream referral sessions and conversions separate: those are observed site outcomes, while response appearance is an off-site observation.

Annotate the timeline with releases, content corrections, major PR, crawler-policy changes, and product-interface changes. If visibility improves after a release, describe the sequence as correlation unless the test design supports a causal conclusion. Generated outputs and retrieval systems change; no honest audit can guarantee persistence.

Leaf’s guide to measuring AI search visibility provides a fuller metric framework. If you need to connect this work with technical SEO, content, structured data, and conversion checks, use the combined SEO and AEO audit guide.

What a finished AI visibility audit delivers

A finished audit is not a gallery of screenshots or a leaderboard without context. It should contain the versioned prompt panel, collection method, response evidence, source inventory, material accuracy findings, explicit limitations, and a prioritized backlog. Every proposed fix should point back to an observation and name how it will be retested.

That standard will not create the comforting certainty of a rank report. It creates something more useful: a reproducible view of where your company is visible, where its claims break down, and which controlled change is worth testing next. Teams can improve the sources and experiences they own. The answer products still control retrieval, generation, presentation, and citation.

Leaf Team
The Leaf team helps businesses and agencies compound organic and AI search traffic. We build the strategy, run the execution, and deliver results — async, systematically, every month.
Back to all posts