Technical SEO
Technical SEO Audit Checklist: 47 Checks in the Order That Matters
A technical SEO audit checklist is a list of pass/fail tests that confirm search engines can crawl, render, index and serve a site’s important pages. The order matters more than the length: a check that fails early in that pipeline makes every check after it meaningless for the pages it affects.
This page is the working list. It has 47 checks grouped as crawl, render, index and serve, one line each, with the test to run, the condition that counts as a pass, and a link to the article that explains the fix. It ends with a script for the first few checks, what the list deliberately leaves out, and answers to common questions.
Key takeaways
- Run the checks in pipeline order. Google processes JavaScript web apps in three phases (crawling, rendering, indexing), and a failure in one phase hides every problem after it.
- Audit templates, not URLs. Pick one representative URL per template, run the list against each, and only then crawl the whole site to size the problems you found.
- Every check needs a pass condition you can test. “Check the sitemap” is a chore; “every sitemap URL returns 200 and is its own canonical” is a check.
- A failing robots.txt fetch is the most expensive single fault. Per Google’s robots.txt specification, Google stops crawling the site for the first 12 hours when robots.txt returns a server error.
- Most checks here are one line because the reasoning lives elsewhere. Each links to the deep-dive article on this site.
Why the order matters
Search engines work through a page in sequence. Googlebot fetches a URL, checks robots.txt, parses the HTML for links, queues the page for rendering and indexes the rendered result. That is how Google’s JavaScript SEO documentation describes it, and it is the order this checklist follows.
The practical consequence is a stop rule. If a template fails a crawl check, do not bother scoring its Core Web Vitals or its schema yet. Fix the upstream fault, then re-run the list from the top for that template.
The fourth group, serve, covers what happens once a page is eligible: how fast it loads for real users, what it says about itself in markup, and whether it stays that way after the next deploy.
How to run it
- List your templates: home, category, product or service, article, tag, search, paginated archive, and any programmatic set.
- Pick one representative URL for each, plus one known-broken URL (a deleted page) and one parameter variant.
- Run the 47 checks against each URL, and record pass, fail or not applicable.
- Crawl the full site with the failures in mind, so the crawl confirms scale rather than generating a 400-row issue export.
That fourth step is where most audits go wrong. A crawler export sorts by issue count, not by pipeline position, so a site with thousands of missing alt attributes and one blocked /products/ directory can end up fixing the alt text first.
Crawl: can the crawler reach the page? (checks 1 to 15)
These checks come first because nothing downstream can happen without a successful fetch.
- robots.txt returns 200. Fetch
/robots.txtdirectly. Pass: a 200 response. Per Google’s robots.txt specification, a 5xx means Google stops crawling the site for the first 12 hours, while most 4xx responses are treated as “no restrictions”. Deep dive: robots.txt for developers. - No priority paths disallowed. Test each template URL against the live file. Pass: no important directory is matched by a
Disallowrule for Googlebot or*. - CSS and JavaScript are crawlable. Test the asset paths each template loads. Pass: none are disallowed, because Google won’t render JavaScript from blocked files.
- robots.txt is valid and under the size limit. Pass: UTF-8 plain text under 500 KiB; the same spec says Google ignores content after the maximum file size.
- AI crawler rules are a decision, not an accident. Pass: each AI user agent is allowed or blocked on purpose. OpenAI’s crawler documentation treats OAI-SearchBot (ChatGPT search) and GPTBot (model training) as independent settings, so blocking one does not block the other. Deep dive: AI crawler log analysis.
- Sitemap is declared and submitted. Pass: a
Sitemap:line in robots.txt with an absolute URL, and the sitemap submitted in Search Console with no processing errors. - Sitemap lists only canonical, indexable 200 URLs. Crawl the sitemap as a list. Pass: zero redirects, 404s, noindexed or non-canonical URLs.
- Sitemap files are within limits. Pass: each file holds at most 50,000 URLs and 50MB uncompressed, the limit in Google’s sitemap guidelines, with larger sets split under a sitemap index.
lastmodis truthful. Pass:lastmodchanges only when main content, structured data or links change. The same Google page says it ignorespriorityandchangefreq, and useslastmodonly if it is consistently and verifiably accurate. Deep dive: crawl budget.- Links are real
<a href>elements. Inspect navigation, pagination and in-content links. Pass: every link to a priority page is an<a>element with anhref, not a click handler or a#fragment route. Deep dive: JavaScript SEO. - No orphaned priority pages. Compare sitemap URLs against the URLs a crawler reaches from the home page. Pass: every priority URL in the sitemap is also linked internally. Deep dive: orphan pages.
- Click depth is shallow for priority templates. Pass: money pages and key articles sit a small number of clicks from the home page, and depth does not grow with every new post. Deep dive: internal linking audit.
- Redirects resolve in one hop. Pass: internal links point at final URLs and no chain has more than one hop. Googlebot follows up to 10 hops by default, according to Google’s HTTP status code documentation, but every hop wastes a fetch. Deep dive: SEO redirects.
- Removed pages return 404 or 410. Request a deleted URL. Pass: a 404 or 410 status, not a 200 error page and not a blanket redirect to the home page.
- Server errors are rare and short. Check logs and the Crawl Stats report. Pass: no sustained 5xx or 429 responses to Googlebot, since both make Google’s crawlers slow down. Deep dive: log file analysis.
If a template fails any of 1 to 15, stop there for that template. The render checks will report on a page Google may never fetch.
Render: does the page contain its content once fetched? (checks 16 to 25)
A successful fetch can still return an empty shell. These checks compare what the server sends with what a renderer builds.
- Primary content is in the initial HTML. Compare
view-source:with the rendered DOM. Pass: headings, body copy and internal links exist in the raw response, not only after JavaScript runs. Deep dive: JavaScript SEO. - Google’s rendered HTML matches the page. Run URL Inspection and open the rendered HTML. Pass: the main content, links and structured data are all present.
- No content waits for interaction. Pass: nothing important loads only on click, swipe or scroll. Google’s mobile-first indexing guide states that Google won’t load content that requires user interactions. Deep dive: pagination and infinite scroll.
- Canonical and robots meta are set in the HTML, not by JavaScript. Pass: the raw HTML carries the final canonical and robots values, and scripts do not change them. Google warns that when it sees
noindexit may skip rendering, so removing it with JavaScript may not work. - Client-side routes use real URLs. Pass: an SPA uses the History API, not
#/fragments, for indexable views. Deep dive: single-page application SEO. - Client-side “not found” views are not soft 404s. Load a missing product or article ID. Pass: the app redirects to a URL that returns 404, or adds
noindex, which are the two options Google documents for client-rendered apps. - Mobile and desktop serve equivalent content. Pass: same primary content, headings, structured data, titles and robots meta on both, because Google indexes the mobile version.
- Static assets are fingerprinted. Pass: JS and CSS filenames change when their content changes (for example
main.2bb85551.js), because Google’s renderer caches aggressively and may ignore caching headers. - Each template uses a rendering mode that fits it. Pass: indexable templates are server-rendered or statically generated; client-side rendering is limited to logged-in or non-indexable views. Deep dive: headless architecture, with framework specifics for React, Next.js and Vue.
- The page works without JavaScript for other bots. Fetch the page with
curl. Pass: the answer to the page’s main question is readable in the response. Google’s own guide notes that not all bots can run JavaScript.
Index: will the engine keep this URL, and which one? (checks 26 to 37)
Crawled and rendered is not the same as indexed. These checks cover the signals that decide whether a URL is kept and which duplicate represents it.
- The Page Indexing report is triaged. Pass: every “not indexed” reason with priority URLs in it has an owner and a cause. Deep dive: crawled, currently not indexed.
- “Discovered, currently not indexed” is small for priority templates. Pass: priority URLs are not waiting in the discovered state for weeks. Deep dive: discovered, currently not indexed.
noindexis only where you meant it. Crawl fornoindexin meta tags andX-Robots-Tagheaders. Pass: every instance is intentional, and none survived a staging-to-production deploy.noindexpages are not also blocked in robots.txt. Pass: a page you want out of the index is crawlable. Google’s noindex documentation says a robots.txt block means the crawler never sees thenoindex, and the URL can still appear in results. Deep dive: canonical vs noindex.- Every indexable page has a self-referencing, absolute canonical in the
<head>. Pass: onerel="canonical", an absolute URL, inside a valid<head>. Relative URLs work but are not recommended. - Google agrees with your canonical. In URL Inspection, compare user-declared and Google-selected canonicals. Pass: they match for each template. Deep dive: alternate page with proper canonical tag.
- Canonical signals agree with each other. Pass: the canonical tag, sitemap entry, internal links and any redirects all name the same URL. Google’s canonicalisation guide explicitly says not to specify different canonicals through different methods.
- Host and protocol variants collapse to one. Request
http://,https://,wwwand non-www. Pass: all permanent-redirect to a single HTTPS origin in one hop. - Parameter duplicates point to the clean URL. Load a page with
?utm_source=and a sort parameter. Pass: both carry the clean URL as canonical, and internal links never use them. - Paginated pages canonicalise to themselves. Pass: page 2 of an archive is its own canonical rather than pointing at page 1, so items beyond page 1 stay discoverable. Deep dive: pagination SEO.
- hreflang is reciprocal and self-referencing. Pass: each language version lists itself and every alternate, with absolute URLs. Google’s localised versions guide says that if two pages don’t both point to each other, the tags are ignored. Deep dive: hreflang and canonical tags.
- hreflang codes are valid and canonicals stay in-language. Pass: ISO 639-1 language codes with optional ISO 3166-1 Alpha 2 regions (
en-GB, neveren-UK), anx-defaultwhere you have a selector page, and each language version canonicalising to a page in the same language. Deep dive: international SEO.
Serve: what happens when the page is shown and used? (checks 38 to 47)
A page that is crawled, rendered and indexed still has to load well, describe itself accurately and stay healthy after the next deploy.
- HTTPS is clean. Pass: a valid certificate for the exact host, no HTTPS-to-HTTP redirects and no insecure dependencies. Google’s canonicalisation guide lists each of these as a reason it may prefer the HTTP version.
- Core Web Vitals pass on field data. Use the Search Console Core Web Vitals report or PageSpeed Insights field data, not a single Lighthouse run. Pass: each priority template is rated good for LCP, INP and CLS. Deep dive: Core Web Vitals for developers.
- The LCP element is discoverable early. Pass: the hero image or heading that becomes LCP is in the initial HTML, is not lazy-loaded, and is not injected by a script after load.
- Layout space is reserved. Pass: images, video and ad or embed slots carry explicit dimensions, and web fonts do not reflow the first screen.
- The server stays fast under crawl load. Pass: response times are stable in the Crawl Stats report. Google’s crawl budget guide says the crawl capacity limit goes down when a site slows or returns server errors.
- Conditional requests are supported. Pass: unchanged pages can answer
If-Modified-Sincewith a304 (Not Modified), which the same guide recommends to save server bandwidth and resources. - Titles and meta descriptions are unique per indexable URL. Pass: no duplicates across templates, both present in the raw HTML, and the title carries the page’s main query. Deep dive: SEO for engineers.
- Structured data is valid and truthful. Pass: JSON-LD is in the server response, validates without errors, and describes only content visible on the page. Google’s structured data guidelines say not to mark up content that is not visible to readers. Deep dive: structured data for AI search.
- Each template carries the schema types that fit it. Pass: Article on articles, BreadcrumbList where breadcrumbs are shown, and one Organization entity, referenced by
@idrather than redefined on every page. - Checks re-run on every deploy. Pass: robots.txt, status codes, canonicals,
noindexand structured data on the template URLs are tested automatically after each release, with an alert when any of them changes. Deep dive: Search Console API.
Core Web Vitals thresholds for check 39
The thresholds are published on web.dev’s Web Vitals page. LCP is good at 2.5s or under and poor above 4.0s; INP is good at 200ms or under, needs improvement between 200 and 500ms, and is poor above 500ms; CLS is good at 0.1 or under and poor above 0.25. Each is assessed at the 75th percentile of page loads, segmented by mobile and desktop, and a page passes only when all three meet the target.
Script the first checks
Several crawl and index checks can be run from a terminal before you open a crawler. This script covers checks 1, 13, 14 (run it on a deleted URL), 19, 28 and 30 for a single URL.
#!/usr/bin/env bash
# Usage: ./tech-audit.sh https://example.com/some-page/
set -u
url="$1"
origin=$(printf '%s' "$url" | sed -E 's#^(https?://[^/]+).*#\1#')
status() { curl -s -o /dev/null -w '%{http_code}' "$1"; }
echo "robots.txt status : $(status "$origin/robots.txt")"
echo "page status : $(status "$url")"
echo "redirect hops : $(curl -s -o /dev/null -L -w '%{num_redirects}' "$url")"
echo "x-robots-tag : $(curl -sI "$url" | grep -i '^x-robots-tag' | tr -d '\r')"
html=$(curl -sL "$url")
echo "canonical : $(printf '%s' "$html" | grep -oiE '<link[^>]+rel="canonical"[^>]*>' | head -1)"
echo "meta robots : $(printf '%s' "$html" | grep -oiE '<meta[^>]+name="robots"[^>]*>' | head -1)"
echo "h1 in raw HTML : $(printf '%s' "$html" | grep -ciE '<h1[ >]')"
An empty x-robots-tag or meta robots line means the header or tag is absent, which is a pass for an indexable page. A 0 on the last line means the heading only exists after JavaScript runs, so go to check 16.
The script reads the raw response only. It will not tell you what Google rendered, which is why checks 17 and 31 still need URL Inspection.
What this checklist leaves out
This is the technical layer only. Content quality, keyword targeting, internal link strategy beyond crawlability, and off-page authority are separate audits.
The 12-phase SEO and GEO audit places this layer as its second phase, after business context and before entity, content and citation work. The AI-search half has its own working list in the GEO audit checklist. The technical SEO audit methodology covers the reasoning, tooling and cadence behind these checks in long form. If you want the whole list run on your site and the fixes shipped as code, that is the scope of technical SEO and GEO consulting.
Frequently asked questions
What should a technical SEO audit checklist include?
It should include checks for crawl access (robots.txt, sitemaps, links, status codes), rendering (content in the initial HTML, JavaScript dependencies), indexing (noindex, canonicals, hreflang) and serving (Core Web Vitals, HTTPS, titles, structured data). Each check needs a testable pass condition. The 47 checks above follow that order.
What order should technical SEO checks be done in?
Crawl first, then render, then index, then serve. That matches how Google processes a page, and it means a failure early in the list explains the symptoms later in it. If a template fails a crawl check, fix it before scoring anything downstream.
How long does a technical SEO audit take?
Running this list against one URL per template takes a few hours for a typical site with ten or so templates. Sizing the problems with a full crawl, reading server logs and writing developer-ready fixes takes longer, and depends on site size and how many templates fail early checks.
How often should you run a technical SEO audit?
Run the full list at least once a year and after any migration, redesign or framework change. Automate the checks that break silently (robots.txt, status codes, canonicals, noindex and structured data) so they run after every deploy.
Can I do a technical SEO audit without paid tools?
Yes, for most of this list. Search Console (URL Inspection, Page Indexing, Core Web Vitals and Crawl Stats), PageSpeed Insights, a browser’s view-source and curl cover the majority of checks. A paid crawler mainly helps with scale: finding every orphan, chain and duplicate across thousands of URLs.
Is crawl budget worth checking on a small site?
Usually not. Google’s crawl budget guide is aimed at sites with over a million unique pages that change weekly, or over 10,000 that change daily, and it describes those numbers as rough estimates. Smaller sites should still check status codes and redirects, because those affect every site.