AEOAI SearchCompetitor AnalysisRecommendations

Why ChatGPT recommends competitors

Audit why competitors enter AI shortlists by testing discovery, inclusion, rank, rationale, caveats, citations, accuracy, and buyer fit separately.

Leaf Team
August 12, 2026
8 min read

ChatGPT may recommend a competitor because that company appears to fit the buyer’s stated constraint better in the sampled answer, or because the available evidence represents it more clearly. The useful response is to reconstruct the shortlist decision, verify its rationale, and repeat the test.

Consider a fully hypothetical procurement case. A 900-person software company asks for a support platform that can keep EU customer data in the EU, connect to Salesforce, and be deployed by a two-person operations team within six weeks. The answer mentions the buyer’s existing vendor, Northstar, but recommends a fictional competitor, Cedar. Its decisive passage reads: “Choose Cedar for this rollout because its documentation explicitly states EU data residency and a managed Salesforce connector. Northstar appears suitable for larger custom implementations, but I could not verify EU residency for this use case.”

That passage gives the audit something precise to test. Did both brands enter the shortlist? Was EU residency truly decisive? Does the cited material support the distinction? Is the statement about Northstar accurate? The investigation stays with this buyer, prompt, sample, and visible evidence. It does not attempt to reconstruct hidden model weights or a universal ranking formula.

Reconstruct the shortlist decision

Read the complete answer before counting names. Record whether each brand is included, absent, or explicitly excluded. Capture position only when the response presents a genuine ranking. If it calls the list “not ranked,” do not convert composition order into preference.

Use one row per brand per run. In the constructed case, the first hypothetical record could look like this:

Field Northstar Cedar
Discovery Mentioned without being named in the prompt Mentioned without being named in the prompt
Shortlist status Included, then not selected Included and recommended
Stated fit Salesforce possible, with implementation framed as heavier Managed Salesforce connector and six-week rollout framed as plausible
Decisive constraint EU residency not verified in the answer EU residency described as explicit
Caveat “Larger custom implementations” No caveat attached to residency
Visible support Current integration page, with no residency passage found Current residency and connector passages shown
Accuracy review Integration description supported, but implementation claim unsupported Quoted capabilities supported in this hypothetical record
Reviewer confidence High on evidence gap and low on implementation claim High on cited capability fit

Keep the raw answer and the exact supporting passages beside this coded record. Also keep referral, conversion, and other business evidence in a separate outcome field. A citation does not establish a visit, and a mention does not establish a qualified opportunity.

This discipline prevents casual mentions from inflating recommendation metrics. Leaf’s guide to measuring AI search visibility defines the relevant denominators. The AI brand perception audit extends the record across access, retrieval, entity, sentiment, and business-value fields.

Test discovery without naming either brand

A prompt such as “Tell me about Northstar” supplies the entity. It tests recognition and factual representation, not spontaneous discovery. Keep those branded checks separate from neutral category and problem prompts.

Build neutral prompts from real buying constraints: organization size, region, budget model, integrations, regulatory needs, implementation capacity, and use case. Include only details that would genuinely affect the decision. Freeze the wording and record the product, visible mode, date, market, account state, and repetition. Compare brands under identical conditions.

The hypothetical case begins with the neutral procurement prompt above. Both brands appear in the answer, so the sampled failure is not discovery. When the same test is repeated in this constructed example, suppose both appear in five of six eligible runs and Cedar is recommended in four. Northstar is therefore being considered, and its gap occurs later in the decision. These are explicitly illustrative results, not observed platform data or a recommended universal sample size.

If a real brand appeared only after being named, the next investigation would concern entity and category discovery. If it appeared neutrally but was excluded after constraints were added, the investigation would move to fit, evidence, or representation. Preserve those stages rather than collapsing them into one visibility score.

Verify the exact rationale

Open every visible citation. Identify the claim it supports, its date, and its source type: owned material, independent coverage, primary research, directory, review platform, or competitor content. A source may support one capability while leaving the comparative conclusion unsubstantiated.

In the constructed record, Cedar’s displayed passages explicitly support EU residency and the managed connector. Northstar’s current integration page supports Salesforce connectivity, but its accessible materials in the sample contain no explicit EU-residency statement. The answer’s claim that Northstar suits “larger custom implementations” has no visible support. That creates two separate findings:

  1. A specific fit-and-evidence gap: for this EU-residency-constrained buyer, Cedar has direct supporting evidence that Northstar’s sampled materials lack.
  2. An accuracy issue: the implementation-size characterization is unsupported and should be logged rather than accepted as explanation.

The first finding earns the recommendation diagnosis in this hypothetical case. It is not merely one item on a list of possible causes. Cedar wins because the decisive buyer requirement is supported for Cedar and unverified for Northstar in the reviewed evidence.

Use the same attribute taxonomy for every brand. “Cedar has more authority” is too vague. “The answer cites an explicit EU-residency passage for Cedar but none for Northstar” can be checked, assigned, and retested. Likewise, do not copy an invented competitor advantage into your site. Track unsupported passages as accuracy incidents.

Google says AI Overviews and AI Mode may use query fan-out, issuing related searches across subtopics and data sources, in its AI features documentation. OpenAI documents crawler and access roles in its official bot documentation. These public details inform collection and access checks without explaining how a specific recommendation was generated.

Classify the gap at the stage where it occurs

Five labels keep unlike failures apart:

A real case can carry more than one label. The constructed Northstar record has a primary fit-and-evidence finding plus a separate factual-representation incident. Both candidates appeared neutrally, and the same rationale recurred in four of the six illustrative runs.

Turn the finding into a controlled action

The action should follow the demonstrated gap. In this hypothetical case, Northstar’s owner first confirms whether the product actually offers EU data residency. If it does not, the exclusion is useful buyer guidance and no visibility fix is appropriate. If it does, the owner publishes or clarifies a dated canonical statement, checks access to that evidence, corrects inconsistent profiles, and schedules the frozen prompt for retesting. The unsupported implementation claim remains a separate accuracy ticket.

Build the backlog with the prompt, exact exclusion or caveat, materiality, evidence, owner, action, and retest date. Prioritize wrong safety claims, namesake confusion, and omitted required capabilities over one unexplained absence in a low-value prompt. Leaf’s AEO assessment can structure checks across technical access, content, authority signals, and measurement before the retest.

Set repetitions before collection according to decision risk and operating capacity. Preserve every eligible run and annotate changes in wording, product mode, or interface. If recommendation inclusion rises after evidence work, report the increase as an observation. Unless the design isolates the intervention, do not claim that the change caused it.

For the hypothetical buyer, the defensible answer is now concise: Cedar was recommended in the sampled case because its visible evidence directly addressed the decisive EU-residency requirement, while Northstar’s did not. One additional statement about Northstar lacked support and requires correction. The result is specific enough to act on within the constructed sample.

Frequently asked questions

How does ChatGPT choose which businesses to recommend?

No public formula explains every recommendation. Test the same buyer prompt repeatedly, save inclusion, rationale, caveats, and citations, and verify claims. The output reveals observable patterns, not internal model weights.

Why does ChatGPT recommend competitors but omit my company?

Possible reasons include buyer fit, entity confusion, source gaps, stale facts, access problems, prompt wording, or normal output variation. Diagnose each with identical prompts, source review, technical checks, and repeated runs before acting.

How can a brand audit its visibility across AI recommendation engines?

Use a fixed, buyer-relevant panel across named products and record full answers, conditions, mentions, shortlist inclusion, rationale, caveats, citations, and accuracy. Report raw counts and product-specific results rather than one universal score.

Does a direct branded prompt inflate apparent recommendation visibility?

Yes, if it is combined with neutral discovery metrics. Naming the brand supplies the entity and tests factual representation. Keep branded checks separate from prompts where the product must discover candidates.

What evidence supports an AI-generated brand recommendation?

Inspect the answer’s rationale and every visible citation, then verify material claims against current primary sources. Citations can support parts of an answer, but they do not expose all inputs or guarantee that the recommendation is correct.

How many times should recommendation prompts be repeated?

Set repetitions before collection based on prompt value and risk, keep the policy stable, and report every run. Increase sampling when answers are volatile or the decision is consequential.

Leaf Team
The Leaf team helps businesses and agencies compound organic and AI search traffic. We build the strategy, run the execution, and deliver results — async, systematically, every month.
Back to all posts