How to audit a website for SEO in 10 steps
A practical 10-step website SEO audit covering business scope, crawling, index controls, rendering, content, internal links, structured data, and measurement.
A website SEO audit should explain which problems limit qualified discovery, which pages and templates are affected, and what the team should fix first. It is not a crawler export with every warning set to high priority. The work combines technical evidence, content judgment, business context, and a clear retest for each recommendation.
The ten-step process below is designed for B2B sites. It starts with revenue exposure, works through access and page quality, and ends with measurement and prioritization. AI-answer observations can be added after the foundations are checked, but they should remain a separate, variable evidence layer.
Step 1: define commercial scope and success
List the offers, audiences, markets, and conversion paths that matter. Identify revenue pages, high-value educational pages, major templates, and business-critical journeys such as demo requests, contact forms, trials, or partner inquiries.
Write down what the audit must help decide. Examples include whether to rebuild a product template, consolidate a resource library, repair an international setup, or improve discovery for a new category. Pull a baseline from analytics and Search Console where access exists, but verify the instrumentation before trusting it.
Scope exclusions matter. Log analysis, backlink review, international targeting, migration planning, and full content inventory may each require additional work. Name what is included rather than implying one crawl covers every discipline.
Step 2: crawl the site and classify URLs
Crawl the public site using the same starting points search crawlers use: internal links and declared sitemaps. Record final status, redirect chain, canonical, robots directives, title, headings, content signals, internal links, structured data, and indexability indicators.
Group URLs by template and business role. Pattern analysis is more useful than treating 4,000 duplicate-title rows as 4,000 independent recommendations. Separate product, solution, article, profile, utility, filter, parameter, and system URLs. Compare the crawler’s discovered set with sitemap URLs, analytics landing pages, and Search Console pages to find orphaned or legacy sections.
A crawl is an observation from one client under one configuration. It does not prove that Google has crawled or indexed the same URLs. Preserve crawl settings and timestamp so the result can be repeated.
Step 3: check access, index controls, and canonicalization
For important URL patterns, inspect response status, redirects, robots.txt, meta robots, X-Robots-Tag, canonical links, sitemap membership, and internal links together. These controls have different jobs.
Google documents that robots.txt manages crawler access, not guaranteed exclusion from search results; a blocked URL can still appear without a snippet if discovered elsewhere (Google Search Central). Use noindex on crawlable pages when indexing exclusion is the goal, following current platform guidance. Do not block the page in robots.txt and then assume a crawler can see its noindex directive.
Check canonical signals for agreement. A canonical hint pointing to one URL while internal links and sitemaps promote another creates avoidable ambiguity. Validate redirect destinations and loops. For JavaScript applications, test the actual HTTP response and rendered result rather than assuming the browser experience proves crawler access.
Step 4: validate rendering and template behavior
Compare raw HTML with rendered DOM on representative templates. Confirm that primary copy, headings, links, metadata, canonical tags, and structured data are present and stable after rendering. Test error, empty, filtered, paginated, and signed-out states where relevant.
JavaScript is not automatically an SEO defect. The audit question is whether important content and links are available reliably to the intended crawler and user. Look for rendering failures, delayed or interaction-only links, client-side soft 404s, metadata overwritten after load, and duplicate routes created by application state.
Use browser inspection and, where appropriate, Google’s URL Inspection tools to examine rendered output. Sample each shared template plus exceptions; one successful homepage render says little about product or documentation routes.
Step 5: review performance with field and lab evidence
Separate real-user field data from controlled lab diagnostics. Google’s Core Web Vitals currently focus on Largest Contentful Paint, Interaction to Next Paint, and Cumulative Layout Shift; web.dev publishes definitions and thresholds (web.dev). Field data describes observed user experiences when enough data is available. Lab tools help reproduce and diagnose specific conditions.
Report the data source, period, device class, URL or origin scope, and sample availability. Do not substitute a single fast desktop test for mobile field evidence. Investigate patterns by template and component: hero media affecting LCP, third-party scripts delaying interaction, or banners causing layout shifts.
Performance recommendations need implementation detail. “Improve Core Web Vitals” is unfinished. “Preload the confirmed hero image on the product template and verify p75 mobile LCP in field data after release” is an actionable test, although field improvement still depends on traffic and collection time.
Step 6: map intent and audit page quality
Assign each important page a primary audience, task, and stage. Check whether the title, H1, opening, sections, and conversion path support that purpose. Identify pages competing for the same intent and important questions with no credible destination.
Review content for directness, original value, evidence, maintenance, and claim accuracy. Google’s people-first content guidance provides useful self-assessment questions, but it is not a checklist that guarantees ranking.
For B2B pages, verify capabilities, pricing context, integrations, proof, limitations, and implementation claims against authoritative sources. Consolidate near-duplicates only after preserving distinct value. A page with low traffic is not automatically useless; it may support a narrow but commercially important decision.
Step 7: inspect internal links and information architecture
Trace how users and crawlers reach revenue pages from the homepage, navigation, topic hubs, related resources, and strong inbound landing pages. Record orphan pages, excessive depth, broken links, generic anchor text, and template links that create large low-value crawl spaces.
Internal links should express useful relationships. Link a definition to the deeper guide, an integration article to the maintained integration page, and a comparison guide to the relevant product route. Avoid inserting repetitive exact-match anchors into unrelated paragraphs.
Use path and inlink data to identify structural fixes, then inspect the rendered component. A link present in raw markup but hidden in an inactive or broken state may not serve users. Conversely, repeated navigation and footer links can inflate raw counts without proving contextual prominence.
Step 8: validate structured data and entity consistency
Inventory JSON-LD by template. Parse each script, validate types and properties, and compare values with visible content. Use the Schema.org validator for vocabulary review and Google’s feature-specific documentation for supported search experiences.
Google states that correct structured data does not guarantee a rich result. Markup should describe the page accurately, not invent ratings, authors, prices, or FAQ content. Keep organization, product, service, and author identities consistent across templates.
See Leaf’s schema markup for AI search guide for a four-layer validation method. Schema can reduce ambiguity and support eligible features; it cannot compensate for inaccessible pages or weak information.
Step 9: verify analytics and search measurement
Test that analytics loads with the intended consent behavior and that key events fire once with correct parameters. Submit test forms through safe staging or designated test paths, then verify events and CRM receipt where access permits. Review channel definitions, referral exclusions, cross-domain behavior, and known data loss.
In Search Console, inspect query and page trends, indexing reports, sitemaps, enhancements, and representative URL inspections. Treat sampled and delayed platform reports according to their documented limits. Annotate migrations, tracking changes, releases, and major campaigns.
If the business cares about AI answer visibility, add a fixed prompt panel with exact products, dates, repetitions, responses, and cited URLs. Leaf’s AI visibility audit explains the protocol. Do not merge sampled answer mentions into organic clicks or call the result market share.
Step 10: prioritize findings and define retests
Score findings using business exposure, severity, confidence, effort, and dependencies. A blocked product template generally comes before rewriting an informational FAQ. A broad claim about “duplicate content” should not outrank a verified conversion failure merely because the crawler produced more rows.
Use this issue format:
| Field | Required detail |
|---|---|
| Observation | What was reproduced, without inferred cause |
| Evidence | Exact URLs, template, export, screenshot, or test |
| Impact | User and business consequence |
| Cause | Confirmed cause or explicitly labeled hypothesis |
| Recommendation | Specific implementation change |
| Owner | Engineering, web, content, analytics, or other |
| Dependency | Work that must happen first |
| Retest | Binary check or monitored experiment |
For example: “Twelve product URLs emit canonical links to the category page” is an observation. The retest is to fetch all twelve after release and confirm self-referencing canonical links plus intended sitemap and internal-link signals. Ranking recovery is a monitored outcome, not the ticket’s pass condition.
A practical website audit checklist
Before delivering the audit, verify that it contains:
- □ Named business goals, critical pages, templates, and exclusions
- □ Reproducible crawl configuration and affected URL exports
- □ Separate evidence for access, index controls, and canonicals
- □ Rendered checks across representative and exceptional states
- □ Field and lab performance data labeled correctly
- □ Intent and claim review for priority pages
- □ Internal-link and architecture findings tied to real journeys
- □ Structured data checked for syntax, meaning, and visible accuracy
- □ Analytics and conversion events tested rather than assumed
- □ Every material issue assigned an owner, dependency, and retest
A checked list does not guarantee rankings. It confirms that the audit has moved from tool output to an implementation-ready diagnosis.
Turn the audit into a controlled improvement cycle
Preserve the baseline, implement high-confidence dependency fixes first, and retest the exact affected set. Monitor search and business outcomes over a suitable period without claiming causation from timing alone. Re-crawl after releases that change routing, templates, or content systems.
If your team needs one workflow spanning technical SEO, content, answer readiness, and visibility sampling, read Leaf’s SEO and AEO audit guide. The central principle is the same: distinguish what you can verify on the site from what an external search or answer system may choose to do. That produces a smaller, clearer backlog—and a much better chance that the audit is actually implemented.