How to Do Technical SEO With AI: My Step-by-Step Workflow (2026) | CrawlRaven
How to Do Technical SEO With AI: My Step-by-Step Workflow (2026)
This workflow leans on one thing: real crawl data for AI to reason over. That's the half we automated — CrawlRaven runs a 200-point technical crawl and hands you a scored, prioritized issue list, so you can spend your AI time on judgment instead of spreadsheet wrangling. Starts at $9/month.
The workflow
My 6-step AI technical SEO loop
Crawl the site
Pull a full technical crawl, then let AI summarize what's broken.Prioritize with AI
Feed the crawl export in and get an impact × effort ranking.Generate fixes
Schema, robots.txt, redirect regex, hreflang — drafted in seconds.Validate everything
Never ship AI output unchecked. Test schema, diff redirects.Make it AI-crawler ready
Allow AI bots, server-render content, confirm it's actually readable.Monitor on a loop
Re-crawl on a schedule so regressions surface fast.
First, get the expectations right
Every disaster I've seen with "AI SEO" comes from the same mistake: treating the model as an oracle instead of an analyst. AI is phenomenal at the boring, high-volume parts of technical SEO and genuinely bad at the parts that decide whether your work matters. Before any prompts, internalize this split:
What AI does well vs. what still needs you
🤖 AI handles the grunt work
- Summarizing 10k-row crawl exports into plain English
- Drafting JSON-LD schema, robots.txt, and redirect regex
- Spotting patterns across huge log files
- Ranking issues by impact × effort in seconds
- Explaining technical findings to non-SEO stakeholders
🧠 You stay in the loop
- Deciding which fixes actually move the business
- Catching AI hallucinations before they ship
- Judging intent, relevance, and brand context
- Verifying schema, redirects, and crawl rules really work
- Owning the call when AI and the data disagree
Rule of thumb: let AI do the reading, sorting, and drafting — then verify every output yourself.
The stack I actually use
You don't need a new platform for this. My entire AI technical SEO setup is four things:
- CrawlRaven's MCP, connected to your LLM — this is the engine.
- A reasoning model — ChatGPT (GPT-5-class) or Claude — for analysis, prioritization, and drafting.
- Google Search Console for real index and performance data the crawler can't see.
- Validators — Google's Rich Results Test, a redirect checker, a robots.txt tester — because nothing AI writes ships unverified.
Step 1: Crawl first, then let AI read the crawl
AI can't audit what it can't see. So I start with real data: a full crawl that gives me status codes, titles, meta, canonicals, indexability, word counts, and response times for every URL.
Step 1 — summarize the crawl
You are a senior technical SEO. I'm attaching a CSV export from a site crawl
(columns: URL, Status Code, Indexability, Title, Meta Description, Canonical,
Word Count, Response Time).
Summarize the technical health of this site in plain English:
1. The 5 most serious issues, by how many URLs each affects.
2. Any patterns (e.g. a section returning 404s, canonical mismatches, thin pages).
3. Anything that looks like it could deindex pages or waste crawl budget.
Be specific and cite example URLs. Don't suggest fixes yet — just diagnose.
Step 2: Make AI prioritize by impact × effort
This is where AI earns its keep. I make the model do that scoring, then sanity-check it. Same export, new prompt:
Step 2 — prioritize the fixes
Using the same crawl data, build a prioritized action plan.
For every distinct issue type, give me a table with:
- Issue
- # of URLs affected
- Impact on rankings/indexing (1–5, with a one-line reason)
- Implementation effort (1–5)
- Quadrant: Quick Win / Major Project / Nice-to-have / Deprioritize
Sort so the highest-impact, lowest-effort items are at the top.
Flag anything that could remove pages from Google's index as CRITICAL, regardless of effort.
Step 3: Let AI draft the tedious fixes
Schema markup, robots.txt rules, redirect regex, hreflang clusters — this is finicky, error-prone, copy-paste work that AI is genuinely great at drafting.
Step 3 — generate JSON-LD schema
Generate valid schema.org JSON-LD for this page. I'll paste the content below.
Requirements:
- Use the most appropriate type(s) (e.g. Article, Product, FAQPage, BreadcrumbList).
- Only include properties you can fill from the content I give you — never invent ratings, prices, or dates.
- Output a single <script type="application/ld+json"> block, ready to paste.
- After the code, list any properties I should add manually and why.
Page content:
"""
[paste the page's visible content, headings, author, publish date here]
"""
Step 4: Validate everything (non-negotiable)
This is the step that separates “AI saved me hours” from “AI deindexed my blog.”
- Schema → paste into Google's Rich Results Test and the Schema validator. If it doesn't validate, it doesn't go live.
- Redirects → run the example URLs through a redirect checker and confirm single-hop 301s, no loops.
- robots.txt → test key URLs in a robots tester so you didn't just block a section you meant to keep.
Step 5: Make the site AI-crawler ready
In 2026, technical SEO isn't just for Googlebot — it's for the AI engines increasingly sending (and answering) queries. Almost none of them render JavaScript.
Step 5 — analyze AI bot traffic in your logs
I'm pasting a sample of my server access logs. Analyze AI/search crawler behavior:
1. List every bot user-agent you see (focus on GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended, Googlebot, Bingbot) and how many requests each made.
2. Which status codes are these bots receiving? Flag any non-200s they hit.
3. Are any of them being served redirects or errors on key pages?
4. Summarize: is anything blocking these bots from reaching my main content?
Logs:
"""
[paste a few hundred log lines]
"""
Step 6: Put it on a loop
A one-time audit is a snapshot; sites rot continuously. The real unlock with AI is that re-running this loop is cheap, so I schedule it instead of waiting for a quarterly panic.
The guardrails that keep this safe
- Real data in, or garbage out. Always feed AI an actual crawl, GSC export, or logs.
- Diagnose, prioritize, and fix in separate prompts. Mixing them lets critical issues hide behind cosmetic ones.
- Validate every output. Schema, redirects, robots — test before deploy, every time.
- Constrain the prompt. Prevent most hallucinations.
- You make the final call. AI ranks and drafts; you decide what matters to the business and own the result.
Key Takeaways
- → AI is the analyst, not the boss: AI removes most of the manual grunt work; it does not replace the SEO.
- → Always start from a real crawl: Feed them an actual export, GSC data, or logs for ground truth.
- → Separate diagnose / prioritize / fix: Three distinct prompts.
- → Validate everything before deploy: Test schema, redirects, robots before it goes live.
- → Most AI bots don't render JS: Server-render anything you want AI search to read and cite.
- → Make it a loop: Schedule a fresh crawl monthly and re-run the summarize-and-prioritize prompts on the new export.