AEOAI SearchTechnical SEOEntity Understanding

Does ChatGPT understand your website?

Test whether ChatGPT can access, retrieve, reconstruct, and accurately cite your website without mistaking a mention for understanding.

Leaf Team
August 12, 2026
6 min read

To test whether ChatGPT understands your website, trace five observable layers: technical access, source retrieval, correct fact extraction, entity reconstruction, and answer citations. Give each layer its own pass criteria and evidence. A mention or citation is one observation within that trace.

“Understand” is shorthand. Define a pass in terms a reviewer can inspect. Can the tested product identify the correct company, reconstruct current material facts, place it in the right category, state relevant limits, and support claims with suitable sources?

Use a website-to-answer trace worksheet

Create one row for each tested fact and prompt:

Layer Evidence to record Next question
Access Status, robots policy, canonical, rendered text, crawler log if available Was the page retained or retrieved?
Retrieval Visible source or citation associated with the answer Which nearby claims does it support?
Fact extraction Exact fact from page compared with answer Are entity relationships also correct?
Entity reconstruction Correct company, product, category, audience, relationships How is the entity framed for this buyer?
Answer Full response, prompt, mode, date, account state Does the result persist across repeated runs?
Citation URL and supported claim Did a measurable site outcome follow?

This trace addresses a gap in many website visibility checks: they treat a mention or citation as proof that an LLM “read” the site. Each layer needs its own evidence.

OpenAI documents named web crawlers and controls in its official bot documentation. Google’s AI features guidance explains technical eligibility while leaving appearance to its systems. Because platform architectures differ, apply each product’s own access rules.

Check access and rendering first

Inspect the canonical pages that should establish company identity, products, audience, evidence, policies, and contact details. Check:

If server logs show a crawler request, preserve the user agent, URL, status, bytes, and time. That is access evidence only. It does not establish that a search index retained the page or an answer retrieved it.

Avoid treating llms.txt, schema, or any single file as a universal access switch. Accurate structured data can clarify visible entities for systems that consume it, but platform eligibility and generated answers remain separate. Leaf’s schema for AI search guide explains this boundary.

Test canonical fact extraction

Build a dated fact sheet from authoritative pages. Include legal and brand names, parent relationships, products, category, audience, regions, current leadership where material, pricing model, key limitations, and discontinued offers. Each fact needs a source URL and review date.

Ask direct questions in fresh, recorded sessions. Examples include “What products does Brand X currently offer?” and “Which customers is Product Y designed for?” Compare each claim with the fact sheet using four labels: correct, incorrect, ambiguous, or unverifiable.

Do not mark a response correct merely because it repeats marketing language. Verify scope. A feature available only on an enterprise plan becomes wrong when presented as universal. A fact can also be current on your site but absent from an answer. That is an omission, not necessarily a false statement.

If errors recur, use Leaf’s protocol for wrong company information in ChatGPT to classify stale sources, collisions, contradictions, and unsupported claims.

Test entity and category reconstruction

Fact recall is narrower than entity understanding. Ask the product to distinguish:

Use neutral prompts as well as branded prompts. A correct answer to “What is Brand X?” demonstrates recognition under a direct cue. Test spontaneous retrieval separately with a question such as “Which providers solve this problem?” The 20-prompt company entity test provides a complete reconstruction panel.

Record uncertainty as a valid outcome. If the answer asks for a domain because the name is ambiguous, that may be safer than confident collision. Do not reward confidence independently of accuracy.

Trace citations and third-party corroboration

For browsing or search-enabled answers, open every citation and map it to the claim it appears to support. Classify the source as owned, primary third-party, independent editorial, directory, review platform, competitor, or unclear.

A third-party citation can corroborate a claim or preserve stale information. Compare its publication date and scope with your canonical source. Request a correction when a publisher’s current page is materially wrong, using its normal editorial process. Do not infer that adding more brand mentions will force an answer change.

Citation review also separates source visibility from sentiment and recommendation. A product may cite your documentation while recommending another brand because of buyer constraints. It may mention you with no owned citation. Business value requires separate analytics evidence and cannot be inferred from citation alone.

Compare products, runs, and changes

Record product, visible mode, account state, market, date, exact prompt, and run number. Repeat important prompts. Do not combine browsing and non-browsing experiences into one result.

After a site update, preserve the old and new page captures, deployment time, crawler observations, and repeated answer results. Keep prompt wording stable. If the answer improves later, report the chronology as correlation unless your design supports a causal claim.

Use a practical pass criterion for each material fact: the correct entity and scoped fact appear consistently enough across the declared repeated sample for the business risk. A safety or legal fact needs a stricter threshold and human escalation than a marketing description.

For a broader technical and content review, Leaf’s SEO and AEO audit can connect this worksheet to indexability, content evidence, schema, measurement, and conversion paths. It cannot guarantee ChatGPT retrieval or citations.

Frequently asked questions

Can ChatGPT pull information from a website?

Yes, some ChatGPT experiences can use web search or links, subject to product behavior and site controls. Record the visible mode and citations. A correct answer alone does not prove that the site was accessed during that run.

How can I test whether an AI crawler can access my website?

Check current official user-agent documentation, robots rules, response status, rendered content, and server logs. A pass means the requested page returned usable content without a relevant block. It does not prove indexing or answer use.

Does a brand mention prove that ChatGPT understands the website?

No. The mention may come from a third-party source, prior data, prompt context, or another path. Test canonical facts, entity relationships, category, limitations, and citations separately.

Which company facts should an LLM reconstruct correctly?

Prioritize facts that identify the entity or change a buyer decision: names, products, category, audience, current capabilities, limits, regions, pricing scope, and material proof. The list should follow business risk, not a universal template.

How do third-party sources affect what ChatGPT understands about a website?

They can corroborate, contextualize, contradict, or preserve stale claims. Capture cited URLs, compare each claim with dated ground truth, and repeat prompts. Visible citations show source presentation, not private system logic.

How should I measure changes after updating the website?

Keep the prompt panel and conditions stable, save before-and-after pages and answers, repeat runs, and report fact accuracy, entity accuracy, mentions, and citations separately. Describe timing without claiming the update caused the answer change.

Leaf Team
The Leaf team helps businesses and agencies compound organic and AI search traffic. We build the strategy, run the execution, and deliver results — async, systematically, every month.
Back to all posts