1001 SEO Media
All posts
By 1001 SEO Mediatechnical SEOaudits

The Technical SEO Audit Checklist We Run on Every New Client

The exact sequence we follow when auditing a site: crawling, indexing, rendering, speed, and structured data, with the traps we see most often.

Every engagement we take on starts with the same technical audit. Not because every site has the same problems, but because the order of investigation matters. There is no point polishing meta titles on pages Google cannot crawl.

The order is the part most teams skip. An audit run as a checklist of equal items produces a list of equal items, which is not a plan. An audit run as a sequence produces a backlog where the first fix unblocks the next measurement. We start at the front door and move inward, because each step depends on the one before it. A site that fails step one will fail every later step in a way that looks like a different problem, and the only way to tell the difference is to fix step one and look again.

Here is the sequence, and the traps we find most often at each step.

1. Can search engines reach your pages?

Start at the front door.

robots.txt: check for accidental Disallow rules, especially ones left over from a staging environment. Also check what you are telling AI crawlers (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot). Many CDN bot-protection presets block them silently. OpenAI documents those user agents in its bot docs.

The staging leftover is the trap we see most. A Disallow rule written to keep a staging site out of Google gets copied to production, and production quietly stops being crawled. The rule is usually one line, and it is usually the last place anyone looks, because the symptom is a ranking drop, not a robots.txt error. The AI crawler side is newer. CDN bot-protection presets grew up to block scrapers, and they block AI crawlers with the same rule. A site that wants to be cited by ChatGPT cannot block the bot that would cite it.

HTTP status codes: crawl the site and look for redirect chains, soft 404s (pages that say "not found" but return 200), and internal links pointing at redirects.

Redirect chains waste crawl budget and dilute signal. A chain of three redirects is three hops the crawler takes before it reaches the destination, and each hop can lose link equity. Soft 404s are worse because they look like real pages to a crawler: a 200 status with "not found" content tells Google the page exists, which pollutes the index with empty pages. Internal links pointing at redirects are a quiet leak, because every link that points at a redirect is a link that does not pass value to its final target.

CDN and WAF rules: firewalls are a common silent killer. If Search Console shows "crawl anomaly" spikes, check your CDN logs before anything else.

On retainers we read AI crawler hits in Promptwatch Agent Analytics on Professional, Business, or an agency plan. Essential has no listed crawler-log allowance. A 403 on a citation candidate is an engineering ticket. It is not a content rewrite.

The distinction between an engineering ticket and a content rewrite is the one that saves budget. A page that returns 403 to ChatGPTBot will never be cited no matter how good the content is. Rewriting the content is the wrong fix. Opening the firewall rule is the right fix. Reading the crawler log first tells you which fix to write, which is why the log is the first thing we open on a retainer.

2. Are the right pages indexed?

Indexed is not the same as crawlable.

Compare the number of pages you want indexed against what site: searches and Search Console's coverage report show.

Hunt for index bloat: faceted navigation, tag archives, internal search results, and URL parameters generating thousands of near-duplicate pages.

Verify canonicals point where you intend. Self-referencing canonicals on parameterized URLs are a classic mistake.

Check pagination and infinite scroll. Content that only loads on scroll may never be seen.

Index bloat is the trap that hides in plain sight. A faceted navigation with three filters can generate thousands of URLs that are near-duplicates of one real page. Each of those URLs consumes crawl budget and dilutes the signal of the canonical. The fix is usually a combination of canonicals, robots rules, and parameter handling, applied to the facet pattern rather than to each URL. Pagination and infinite scroll are the other quiet failure: content that only loads on scroll may never be rendered by a crawler, which means it exists for users but not for the index.

3. Does your content survive rendering?

Modern sites often serve an empty shell and hydrate with JavaScript.

Fetch key pages with JavaScript disabled and compare against the rendered version. Anything critical missing from the raw HTML is at risk.

Watch for client-side-only content: reviews loaded from an API, tabs that inject content on click, and lazy-loaded text below the fold.

If you are on a JS framework, confirm your rendering strategy (SSR, SSG, ISR) actually applies to the templates that matter.

The rendering test is the one that catches the prettiest sites. A site that looks complete in a browser may serve an empty shell to a crawler, with the content hydrated by JavaScript the crawler does not run. Reviews loaded from an API, tabs that inject content on click, and lazy-loaded text below the fold are the three places this fails. The test is simple: fetch the page with JavaScript disabled and read what comes back. If the answer is an empty shell, the content is at risk regardless of how it looks rendered.

4. Is the site fast where it counts?

Core Web Vitals are a ranking signal, but more importantly they are a proxy for user experience.

Prioritize field data (what real users experience) over lab scores.

The usual suspects: oversized hero images, render-blocking third-party scripts, layout shift from ads and late-loading fonts, and slow server response on uncached pages.

Fix templates, not pages. One fix to a product-page template moves thousands of URLs.

Field data over lab scores is the rule that changes the fix list. A lab score is a synthetic run on a fast connection. Field data is what real users on real devices experience. The two often disagree, and the field data is the one that maps to how the page actually performs. Fixing templates instead of pages is the multiplier: one fix to a product-page template moves every product page at once, which is why a template fix beats a page fix every time.

5. Is your structured data earning you anything?

Validate what is there. Broken schema is worse than none.

Add what is missing for your business model: Product, FAQPage, HowTo, LocalBusiness, Article, Organization.

Structured data pulls double duty now: rich results in Google, and cleaner entity extraction for AI engines. Google's AI features documentation is the official page for how Overviews and related features appear. It does not replace a crawl.

Broken schema is worse than none because a broken schema can suppress a rich result that the page would otherwise earn. Validation is the first pass. The second pass is to add the types that match the business model, which is where most sites under-invest. The double duty point is the one that grew in the last year: structured data now feeds both Google rich results and AI entity extraction, which means a schema fix pays twice.

6. Architecture and internal linking

Important pages should be reachable within three clicks of the homepage.

Look for orphan pages (in the sitemap but linked from nowhere). They are common after redesigns and migrations.

Anchor text should describe the target page. "Click here" wastes signal.

Orphan pages are the redesign scar. A migration moves pages into a new structure and forgets to link them from anywhere, but leaves them in the sitemap. They sit in the index with no internal authority, which means they rank on external links alone, if they have any. The three-click rule is the test: if an important page is more than three clicks from the homepage, it is not important to the site architecture, regardless of how important it is to the business.

What comes out the other side

The deliverable is not a 60-page PDF of screenshots. It is a prioritized backlog: each issue scored by impact and effort, assigned an owner, and sequenced so quick wins land while bigger fixes are in progress.

An audit that does not turn into shipped changes is a very expensive bookmark. Whether you run this checklist yourself or bring us in to do it, make sure someone owns the follow-through.

The order is the point. Crawl access comes before indexing because a page Google cannot fetch never enters the index conversation. Indexing comes before rendering because a correctly indexed shell with no hydrated content still fails the user. Rendering comes before speed because a fast empty page is still empty. Speed and structured data come before architecture because the foundation has to hold before internal linking can route authority. Skipping ahead to titles and meta is the most common waste of an audit budget, and it is the step that matters least when the steps above it are broken.

FAQ

Why start with crawl access instead of titles?

There is no point polishing meta titles on pages Google cannot crawl. robots.txt leftovers, WAF rules, and blocked AI crawlers are cheaper to find than a rewrite.

Do we treat GPTBot the same as Googlebot?

No. Many CDN bot-protection presets block GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot silently. OpenAI documents those user agents. We check logs, not just robots.txt intentions.

What do we deliver?

A prioritized backlog, not a 60-page PDF of screenshots. Each issue scored by impact and effort, assigned an owner, and sequenced so quick wins land while bigger fixes are in progress.

If you want this checklist run on a site, hello@1001seomedia.com.