Free AEO audit tools: what they can actually tell you
A candid guide to free AEO audit tools: which checks are reliable, what scores conceal, how to verify findings, and when a broader audit is justified.
A free AEO audit tool is useful when it shows you an observable issue and the evidence behind it. It may confirm that a homepage is reachable, detect a missing canonical, parse visible headings, inventory structured data, find discovery files, or sample how a few prompts are answered. It cannot prove why an answer system omitted your company, predict future citations, or replace business and technical judgment across a full site.
The right way to use a free tool is as triage. Let it identify a narrow question worth investigating. Then verify important findings manually before asking developers, writers, or executives to act.
Know what kind of “free audit” you are running
Tools that share the same label may perform entirely different work. Before interpreting a score, identify the product type:
| Tool type | Typical input | Useful output | Important limit |
|---|---|---|---|
| Single-page checker | One public URL | Metadata, headings, links, markup, response observations | Does not represent the whole site |
| Crawler | Domain or URL list | Status codes, canonicals, directives, internal-link patterns | Business meaning and rendered states may need review |
| Structured-data validator | URL or code | Syntax, detected types, feature eligibility feedback | Valid markup can still be inaccurate or irrelevant |
| Performance test | URL | Lab diagnostics or available field data | Performance is one part of search and conversion quality |
| Prompt monitor | Brand, topics, prompts | Sampled mentions, citations, competitors, answer text | Outputs vary and prompt selection shapes the result |
| Search-console interface | Verified property | Search performance and index-related reporting | Historical platform data, not an AEO diagnosis |
A tool can combine several types, but it should say which checks it actually runs. “AI-ready” is not a test specification.
Leaf’s free SEO and AEO assessment, for example, is a bounded public-homepage review. It checks observable technical, content, structured-data, trust, and discovery signals. It does not log in, execute the site’s scripts, measure rankings or backlinks, or determine whether an AI system will cite the domain. Those exclusions are part of the result, not fine print to ignore.
Deterministic checks are where automation is strongest
Automated tools are most reliable when the input and pass condition are explicit. A URL either returned a recorded status to that request or it did not. A canonical element was present in the retrieved document or it was not. A JSON-LD block parsed under a stated parser or produced an error.
Even these checks need context. A tool’s fetch can differ from a search crawler’s fetch. JavaScript may add or alter content after the initial response. A canonical can be syntactically valid and strategically wrong. A robots rule can block one user agent while allowing another.
Google’s documentation explains that robots.txt controls crawling rather than guaranteed indexing. If a tool reports “blocked from Google” or “deindexed” based only on one robots rule, inspect the exact user agent, path, meta directives, canonical, links, and platform reporting before accepting the conclusion.
Structured data needs similar restraint. Google provides an official Rich Results Test and structured-data guidance, but valid markup does not guarantee a rich result. It also does not prove the facts are true. Compare markup with visible page content and the relevant vocabulary rather than chasing a green icon.
Content and trust checks are indicators, not verdicts
A free tool may flag a missing H1, an unclear opening paragraph, few external sources, no author, or no visible contact link. These are observable features, but their meaning depends on the page.
A product login page does not need an essay. A legal or medical advisory article may need visible authorship, review information, and careful primary sourcing. A pricing page may need clear qualification rules more urgently than additional prose. An about page should support entity and trust questions but does not have to answer every category query.
Be cautious when tools turn page features into universal rules:
- A question heading is not automatically a better answer.
- More words do not automatically create more authority.
- More outbound links do not automatically improve source quality.
- An author box does not prove expertise.
- Schema does not repair unsupported content.
- An llms.txt file is not a universal access or citation control.
Google states that no special AI text file or special schema is required for its AI search features. That is a verified statement about Google’s features, not every answer product. A checker should keep platform-specific facts separate from general recommendations.
Prompt checks require a disclosed sampling method
A tool that tests generated answers can reveal useful observations: whether the brand appeared, which pages were cited, which competitors appeared, and whether the response repeated a wrong fact. The observation becomes difficult to interpret when the prompt set is hidden or designed to favor the brand.
Ask five questions about any prompt-based score:
- What exact prompts were used?
- Which product, model label, market, language, and date were recorded?
- Were prompts neutral or did they name the brand?
- How many times was each prompt run?
- Are mention, citation, recommendation, and factual accuracy separate fields?
OpenAI’s official crawler documentation describes different user agents and controls. Allowing a crawler is an access decision; it is not proof of retrieval, training use, search inclusion, or citation. A tool that collapses all of those into “ChatGPT readiness” is overstating what the check establishes.
Generated answers vary over time and across contexts. Save the raw response and date. Treat a single run as a snapshot. If a decision matters, repeat a fixed prompt under documented conditions and report the denominator.
A composite score can hide the most important failure
Scores are convenient, but weights are editorial choices. A 78/100 may combine dozens of cosmetic passes with one critical broken contact route. Another tool may assign 20 points to optional markup and five points to crawl access. Without the formula, the number is not comparable.
Always open the check-level results. Classify findings into five buckets:
- Access: request, response, directives, rendering, canonical, discovery.
- Content: page purpose, answer clarity, intent fit, internal links.
- Evidence: claims, sources, authorship, dates, company facts.
- Sampled visibility: prompts, mentions, citations, factual errors.
- Conversion: working next step, measurement, qualified action.
Then identify any critical failure regardless of score. An inaccessible pricing page is not offset by correct Open Graph tags. The number can help summarize a bounded method; it should not overrule the evidence.
Use this verification worksheet on any free result
The following artifact turns a generated report into a small, defensible work queue. Complete it for each high- or medium-priority finding before implementation.
| Finding | Tool’s evidence | Manually reproduced? | Fact or heuristic? | Business relevance | Next action |
|---|---|---|---|---|---|
| Example: canonical points elsewhere | HTML element and URL | Yes, on rendered and raw page | Verifiable fact | High on pricing page | Confirm intended canonical with owner |
| Example: “answer too short” | Proprietary score | No objective pass condition | Heuristic | Unknown | Review buyer question; do not pad blindly |
| Example: brand absent from prompt | Saved response and prompt | Yes, one dated run | Sampled observation | Medium if prompt reflects buyers | Repeat fixed panel and inspect cited sources |
| Example: invalid JSON-LD | Parser error and code location | Yes | Verifiable syntax defect | Depends on page and type | Correct code, then validate visible accuracy |
Use these decision rules:
- Act now when a material defect is reproducible, the intended state is known, and the change is low risk.
- Investigate when the symptom is real but the cause or desired state is unclear.
- Experiment when the recommendation aims to influence an external ranking or answer outcome.
- Ignore when the tool applies an irrelevant universal rule or cannot expose evidence.
This is a buyer workflow, not a search-platform standard. Its purpose is to prevent an opaque score from becoming an unreviewed backlog.
Combine free primary-source tools deliberately
No single free tool needs to answer every question. A practical first pass can combine:
- browser and command-line inspection for response, HTML, and rendered behavior;
- the official Google Search Console reports for a verified property’s Google search and indexing observations;
- Google’s structured-data testing resources for eligible markup;
- field and lab performance information described in the official Core Web Vitals guidance;
- a small, versioned prompt worksheet for answer observations;
- analytics and CRM checks for the business action that actually matters.
Label what each source proves. Search Console data can show Google search impressions and clicks reported for the property; it does not reveal every answer engine. A lab performance test can reproduce diagnostic conditions; it does not describe every real visit. A prompt worksheet captures outputs; it does not reveal the hidden cause of those outputs.
Know when free triage has reached its limit
Escalate to a broader audit when the problem crosses templates, systems, or owners. Common triggers include:
- conflicting robots, canonical, rendering, and sitemap signals;
- hundreds or thousands of URLs requiring pattern analysis;
- a migration, redesign, international rollout, or major product repositioning;
- inconsistent product, location, security, or credential claims;
- unclear ownership between marketing pages and documentation;
- a need to connect findings with analytics and qualified pipeline;
- repeated prompt observations that require source and competitor analysis;
- an internal team that needs implementation-ready tickets and retests.
Before paying, compare the exact deliverables using Leaf’s SEO audit deliverables checklist. A professional audit should add scope, context, prioritization, ownership, and verification—not merely rerun the same checker with a larger PDF.
Leaf’s $875 combined SEO and AEO audit is one fixed-scope option for established B2B sites. It is appropriate only when its stated boundaries fit the site and the team can implement the resulting backlog. The free assessment remains the better starting point when you need a quick public-homepage check and understand its limits.
The most valuable free-tool result is rarely the final score. It is a reproducible observation that helps you choose the next sensible investigation without pretending that an external answer system has disclosed its decision process.