← Field notes

Case note: the technical SEO audit that ran itself overnight

Key takeaways

  • A 4,200-page technical SEO audit ran end-to-end in a single overnight pipeline: crawl, classify issues by severity, prioritize by traffic impact, draft fix recommendations — 47 minutes of human review time.
  • The pipeline found 312 issues; a priority model ranked them by estimated traffic impact and flagged 22 as critical — all 22 matched what a senior SEO auditor would have surfaced first.
  • The draft fix recommendations were accepted without edit in 78% of cases; 22% needed revision, mostly for client-specific context the pipeline had no access to.
  • Total client cost: one-fifth of a comparable manual audit, delivered in 18 hours vs. three weeks, with zero back-and-forth on issue prioritization.

A B2B SaaS client came in with 4,200 indexed pages and 14 months since their last full technical audit. The site had grown through two product launches and an acquisition. Nobody had looked at the crawl data since.

The pipeline

Step 1: Full site crawl. A Screaming Frog-equivalent crawl ran via API — every URL collected, normalised, and deduplicated. 4,200 indexed pages; 6,100 total URLs discovered including canonicals, redirects, and pagination.

Step 2: Issue classifier. An agent categorised every URL across five buckets: broken links, missing meta tags, thin content (under 300 words on indexable pages), duplicate title tags, and pages flagging slow Core Web Vitals. Each URL could carry multiple issue types.

Step 3: Priority ranker. A second agent weighted each issue by estimated organic traffic for that URL — pulled from a Semrush export the client provided. Output: a ranked issue list with severity scores from 1–10. The top of the list was not sorted by issue count. It was sorted by traffic at risk.

Step 4: Fix-recommendation drafting. One recommendation block per issue category, written for developer handoff. Specific: which tag, which template, what the fix looks like in code. Not "add a meta description" — "the following 47 product pages share a dynamically generated title tag that duplicates the H1; update the template to append the brand modifier."

Step 5: Human review. A senior SEO reviewed only the top 22 critical items. Total human time: 47 minutes.

The numbers

312 issues found across the full crawl. 22 flagged critical by the priority ranker. All 22 matched what a senior auditor would have surfaced first — no false positives in the critical tier. 78% of the fix recommendations were accepted without edit and passed directly to the dev team. 22% needed revision before handoff.

What the 22% tells you

The revisions were not random. They clustered in three places.

First: CMS constraints the pipeline had no visibility into. The recommendation said "update the canonical tag on these 8 pages" — correct in principle, but the client's CMS generates canonicals automatically and the field is locked. The fix required a different approach entirely.

Second: redirect history for retired product pages. The pipeline flagged broken internal links pointing to three deprecated URLs. What it couldn't know was that those pages had been intentionally delisted after a product sunset, and the redirect strategy was still under legal review.

Third: anything requiring the client's analytics data. Several thin-content flags were technically accurate but commercially wrong — those pages convert at 4x the site average on paid traffic. Without GA4 data in the pipeline, the ranker had no way to know.

These are not edge cases. They are a structural limit of any pipeline that runs without access to internal systems. The 22% is not a failure rate — it is the cost of operating on crawl data alone.

What the pipeline cannot do

Judgment calls about brand tone in meta rewrites. The pipeline drafts recommendations; it does not write final copy for a brand with a specific voice.

Redirect strategy for retired product pages. That requires knowing why a page was retired, what the commercial intent was, and whether the URL carries any link equity worth preserving.

Anything requiring GA4 or Google Search Console data the client hasn't piped in. Traffic estimates from third-party tools are proxies. Priority ranking improves significantly when real click and impression data is available.

The pipeline is a force multiplier on the audit phase. It is not a replacement for the strategist who knows the client's roadmap.

The outcome

The client moved from annual audits to quarterly pipeline runs.

Frequently asked questions

Can AI run a full technical SEO audit without human involvement?

Mostly yes for the crawl, classify, and prioritize stages. In practice, an AI agent can surface and rank the full issue list without human input. Human review is still required for the top critical items — in this case, 22 issues that took 47 minutes to triage. Anything requiring internal system access (CMS data, analytics history, redirect logs) cannot be automated without a direct integration.

How do you prioritize SEO issues with an AI agent?

Each issue is scored by estimated organic traffic impact on the affected URL. Severity is weighted alongside traffic — a broken canonical on a high-traffic page ranks above thin content on a low-traffic one. The output is a ranked list; the top N items go to human review. The model has no visibility into business priority, so human override is always available and expected.

What does an automated SEO audit pipeline cost compared to a manual one?

In this case, one-fifth of a comparable manual audit, delivered in 18 hours versus three weeks. The cost gap widens as site size grows — the pipeline's marginal cost per additional URL is near zero. Manual audits scale linearly with page count; the automated pipeline does not.

Book a 30-min discovery →