Technical SEO Audit
Categories
1. Crawlability
- robots.txt: exists, valid, not blocking important resources
- XML sitemap: exists, referenced in robots.txt, valid format
- Noindex tags: intentional vs accidental
- Crawl depth: important pages within 3 clicks of homepage
- JavaScript rendering: check if critical content requires JS execution
- Crawl budget: for large sites (>10k pages), efficiency matters
AI Crawler Management
As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.
Known AI crawlers:
| Crawler |
Company |
robots.txt token |
Purpose |
| GPTBot |
OpenAI |
GPTBot |
Model training |
| ChatGPT-User |
OpenAI |
ChatGPT-User |
Real-time browsing |
| ClaudeBot |
Anthropic |
ClaudeBot |
Model training |
| PerplexityBot |
Perplexity |
PerplexityBot |
Search index + training |
| Bytespider |
ByteDance |
Bytespider |
Model training |
| Google-Extended |
Google |
Google-Extended |
Gemini training (NOT search) |
| CCBot |
Common Crawl |
CCBot |
Open dataset |
Key distinctions:
- Blocking
Google-Extended prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use Googlebot)
- Blocking
GPTBot prevents OpenAI training but does NOT prevent ChatGPT from citing your content via browsing (ChatGPT-User)
- ~3-5% of websites now use AI-specific robots.txt rules
Example, selective AI crawler blocking:
# Allow search indexing, block AI training crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
# Allow all other crawlers (including Googlebot for search)
User-agent: *
Allow: /
Recommendation: Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the seo-geo skill for full AI visibility optimization.
2. Indexability
- Canonical tags: self-referencing, no conflicts with noindex
- Duplicate content: near-duplicates, parameter URLs, www vs non-www
- Thin content: pages below minimum word counts per type
- Pagination: rel=next/prev or load-more pattern
- Hreflang: correct for multi-language/multi-region sites
- Index bloat: unnecessary pages consuming crawl budget
3. Security
- HTTPS: enforced, valid SSL certificate, no mixed content
- Security headers:
- Content-Security-Policy (CSP)
- Strict-Transport-Security (HSTS)
- X-Frame-Options
- X-Content-Type-Options
- Referrer-Policy
- HSTS preload: check preload list inclusion for high-security sites
4. URL Structure
- Clean URLs: descriptive, hyphenated, no query parameters for content
- Hierarchy: logical folder structure reflecting site architecture
- Redirects: no chains (max 1 hop), 301 for permanent moves
- URL length: flag >100 characters
- Trailing slashes: consistent usage
5. Mobile Optimization
- Responsive design: viewport meta tag, responsive CSS
- Touch targets: minimum 48x48px with 8px spacing
- Font size: minimum 16px base
- No horizontal scroll
- Mobile-first indexing: Google indexes mobile version. Mobile-first indexing is 100% complete as of July 5, 2024. Google now crawls and indexes ALL websites exclusively with the mobile Googlebot user-agent.
6. Core Web Vitals
- LCP (Largest Contentful Paint): target <2.5s
- INP (Interaction to Next Paint): target <200ms
- INP replaced FID on March 12, 2024. FID was fully removed from all Chrome tools (CrUX API, PageSpeed Insights, Lighthouse) on September 9, 2024. Do NOT reference FID anywhere.
- CLS (Cumulative Layout Shift): target <0.1
- Evaluation uses 75th percentile of real user data
- Use PageSpeed Insights API or CrUX data if MCP available
7. Structured Data
- Detection: JSON-LD (preferred), Microdata, RDFa
- Validation against Google's supported types
- See seo-schema skill for full analysis
8. JavaScript Rendering
- Check if content visible in initial HTML vs requires JS
- Identify client-side rendered (CSR) vs server-side rendered (SSR)
- Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues
- Verify dynamic rendering setup if applicable
JavaScript SEO: Canonical & Indexing Guidance (December 2025)
Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:
- Canonical conflicts: If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.
- noindex with JavaScript: If raw HTML contains
<meta name="robots" content="noindex"> but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.
- Non-200 status codes: Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.
- Structured data in JavaScript: Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.
Best practice: Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.
9. IndexNow Protocol
- Check if site supports IndexNow for Bing, Yandex, Naver
- Supported by search engines other than Google
- Recommend implementation for faster indexing on non-Google engines
Agent-Friendly Pages (forward-looking)
AI agents (not just AI summarizers) increasingly read sites through three
channels: vision models on screenshots, raw HTML/DOM, and the accessibility
tree (the cleanest signal). Audit criteria — semantic HTML (real <button>
and <a>, not <div onclick>), label associations, interactive target sizing,
layout stability across templates, cursor: pointer correctness — live in
references/agent-friendly-pages.md.
Audit command
# Render with Playwright + capture accessibility tree, then score
python3 scripts/agent_ux_check.py https://example.com --json
The scanner outputs an Agent-UX score (0-100) plus itemized issues:
- HTML findings: real buttons / anchors,
<div onclick> widgets, semantic
landmarks, inputs without <label for>, inputs without ARIA labels
- Accessibility tree findings: total nodes, interactive nodes, unnamed
interactive elements,
role="generic" ratio
The accessibility-tree snapshot uses Playwright's
page.accessibility.snapshot(interesting_only=False). To capture the tree
without scoring, use python3 scripts/render_page.py <url> --a11y-tree --json.
Surface findings as opportunities, not failures. The standards (WebMCP,
agent UX heuristics) are early — don't gate audits on a sub-100 score.
Output
Technical Score: XX/100
Category Breakdown
| Category |
Status |
Score |
| Crawlability |
pass/warn/fail |
XX/100 |
| Indexability |
pass/warn/fail |
XX/100 |
| Security |
pass/warn/fail |
XX/100 |
| URL Structure |
pass/warn/fail |
XX/100 |
| Mobile |
pass/warn/fail |
XX/100 |
| Core Web Vitals |
pass/warn/fail |
XX/100 |
| Structured Data |
pass/warn/fail |
XX/100 |
| JS Rendering |
pass/warn/fail |
XX/100 |
| IndexNow |
pass/warn/fail |
XX/100 |
Critical Issues (fix immediately)
High Priority (fix within 1 week)
Medium Priority (fix within 1 month)
Low Priority (backlog)
DataForSEO Integration (Optional)
If DataForSEO MCP tools are available, use on_page_instant_pages for real page analysis (status codes, page timing, broken links, on-page checks), on_page_lighthouse for Lighthouse audits (performance, accessibility, SEO scores), and domain_analytics_technologies_domain_technologies for technology stack detection.
Google API Integration (Optional)
If Google API credentials are configured, use python3 scripts/pagespeed_check.py <url> --json for real PSI + CrUX field data (replaces lab-only CWV estimates), python3 scripts/crux_history.py <url> --json for 25-week CWV trends, and python3 scripts/gsc_inspect.py <url> --json for real indexation status per URL.
Error Handling
| Scenario |
Action |
| URL unreachable |
Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. |
| robots.txt not found |
Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. |
| HTTPS not configured |
Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. |
| Core Web Vitals data unavailable |
Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. |
1---2name: seo-technical3description: Technical SEO audit across 9 categories: crawlability, indexability, security, URL structure, mobile, Core Web Vitals, structured data, JavaScript rendering, and IndexNow protocol. Use when user says "technical SEO", "crawl issues", "robots.txt", "Core Web Vitals", "site speed", or "security headers".4license: MIT5---67# Technical SEO Audit89## Categories1011### 1. Crawlability12- robots.txt: exists, valid, not blocking important resources13- XML sitemap: exists, referenced in robots.txt, valid format14- Noindex tags: intentional vs accidental15- Crawl depth: important pages within 3 clicks of homepage16- JavaScript rendering: check if critical content requires JS execution17- Crawl budget: for large sites (>10k pages), efficiency matters1819#### AI Crawler Management2021As of 2025-2026, AI companies actively crawl the web to train models and power AI search. Managing these crawlers via robots.txt is a critical technical SEO consideration.2223**Known AI crawlers:**2425| Crawler | Company | robots.txt token | Purpose |26|---------|---------|-----------------|---------|27| GPTBot | OpenAI | `GPTBot` | Model training |28| ChatGPT-User | OpenAI | `ChatGPT-User` | Real-time browsing |29| ClaudeBot | Anthropic | `ClaudeBot` | Model training |30| PerplexityBot | Perplexity | `PerplexityBot` | Search index + training |31| Bytespider | ByteDance | `Bytespider` | Model training |32| Google-Extended | Google | `Google-Extended` | Gemini training (NOT search) |33| CCBot | Common Crawl | `CCBot` | Open dataset |3435**Key distinctions:**36- Blocking `Google-Extended` prevents Gemini training use but does NOT affect Google Search indexing or AI Overviews (those use `Googlebot`)37- Blocking `GPTBot` prevents OpenAI training but does NOT prevent ChatGPT from citing your content via browsing (`ChatGPT-User`)38- ~3-5% of websites now use AI-specific robots.txt rules3940**Example, selective AI crawler blocking:**41```42# Allow search indexing, block AI training crawlers43User-agent: GPTBot44Disallow: /4546User-agent: Google-Extended47Disallow: /4849User-agent: Bytespider50Disallow: /5152# Allow all other crawlers (including Googlebot for search)53User-agent: *54Allow: /55```5657**Recommendation:** Consider your AI visibility strategy before blocking. Being cited by AI systems drives brand awareness and referral traffic. Cross-reference the `seo-geo` skill for full AI visibility optimization.5859### 2. Indexability60- Canonical tags: self-referencing, no conflicts with noindex61- Duplicate content: near-duplicates, parameter URLs, www vs non-www62- Thin content: pages below minimum word counts per type63- Pagination: rel=next/prev or load-more pattern64- Hreflang: correct for multi-language/multi-region sites65- Index bloat: unnecessary pages consuming crawl budget6667### 3. Security68- HTTPS: enforced, valid SSL certificate, no mixed content69- Security headers:70 - Content-Security-Policy (CSP)71 - Strict-Transport-Security (HSTS)72 - X-Frame-Options73 - X-Content-Type-Options74 - Referrer-Policy75- HSTS preload: check preload list inclusion for high-security sites7677### 4. URL Structure78- Clean URLs: descriptive, hyphenated, no query parameters for content79- Hierarchy: logical folder structure reflecting site architecture80- Redirects: no chains (max 1 hop), 301 for permanent moves81- URL length: flag >100 characters82- Trailing slashes: consistent usage8384### 5. Mobile Optimization85- Responsive design: viewport meta tag, responsive CSS86- Touch targets: minimum 48x48px with 8px spacing87- Font size: minimum 16px base88- No horizontal scroll89- Mobile-first indexing: Google indexes mobile version. **Mobile-first indexing is 100% complete as of July 5, 2024.** Google now crawls and indexes ALL websites exclusively with the mobile Googlebot user-agent.9091### 6. Core Web Vitals92- **LCP** (Largest Contentful Paint): target <2.5s93- **INP** (Interaction to Next Paint): target <200ms94 - INP replaced FID on March 12, 2024. FID was fully removed from all Chrome tools (CrUX API, PageSpeed Insights, Lighthouse) on September 9, 2024. Do NOT reference FID anywhere.95- **CLS** (Cumulative Layout Shift): target <0.196- Evaluation uses 75th percentile of real user data97- Use PageSpeed Insights API or CrUX data if MCP available9899### 7. Structured Data100- Detection: JSON-LD (preferred), Microdata, RDFa101- Validation against Google's supported types102- See seo-schema skill for full analysis103104### 8. JavaScript Rendering105- Check if content visible in initial HTML vs requires JS106- Identify client-side rendered (CSR) vs server-side rendered (SSR)107- Flag SPA frameworks (React, Vue, Angular) that may cause indexing issues108- Verify dynamic rendering setup if applicable109110#### JavaScript SEO: Canonical & Indexing Guidance (December 2025)111112Google updated its JavaScript SEO documentation in December 2025 with critical clarifications:1131141. **Canonical conflicts:** If a canonical tag in raw HTML differs from one injected by JavaScript, Google may use EITHER one. Ensure canonical tags are identical between server-rendered HTML and JS-rendered output.1152. **noindex with JavaScript:** If raw HTML contains `<meta name="robots" content="noindex">` but JavaScript removes it, Google MAY still honor the noindex from raw HTML. Serve correct robots directives in the initial HTML response.1163. **Non-200 status codes:** Google does NOT render JavaScript on pages returning non-200 HTTP status codes. Any content or meta tags injected via JS on error pages will be invisible to Googlebot.1174. **Structured data in JavaScript:** Product, Article, and other structured data injected via JS may face delayed processing. For time-sensitive structured data (especially e-commerce Product markup), include it in the initial server-rendered HTML.118119**Best practice:** Serve critical SEO elements (canonical, meta robots, structured data, title, meta description) in the initial server-rendered HTML rather than relying on JavaScript injection.120121### 9. IndexNow Protocol122- Check if site supports IndexNow for Bing, Yandex, Naver123- Supported by search engines other than Google124- Recommend implementation for faster indexing on non-Google engines125126## Agent-Friendly Pages (forward-looking)127128AI agents (not just AI summarizers) increasingly read sites through three129channels: vision models on screenshots, raw HTML/DOM, and the **accessibility130tree** (the cleanest signal). Audit criteria — semantic HTML (real `<button>`131and `<a>`, not `<div onclick>`), label associations, interactive target sizing,132layout stability across templates, `cursor: pointer` correctness — live in133`references/agent-friendly-pages.md`.134135### Audit command136137```bash138# Render with Playwright + capture accessibility tree, then score139python3 scripts/agent_ux_check.py https://example.com --json140```141142The scanner outputs an Agent-UX score (0-100) plus itemized issues:143- HTML findings: real buttons / anchors, `<div onclick>` widgets, semantic144 landmarks, inputs without `<label for>`, inputs without ARIA labels145- Accessibility tree findings: total nodes, interactive nodes, unnamed146 interactive elements, `role="generic"` ratio147148The accessibility-tree snapshot uses Playwright's149`page.accessibility.snapshot(interesting_only=False)`. To capture the tree150without scoring, use `python3 scripts/render_page.py <url> --a11y-tree --json`.151152Surface findings as **opportunities**, not failures. The standards (WebMCP,153agent UX heuristics) are early — don't gate audits on a sub-100 score.154155## Output156157### Technical Score: XX/100158159### Category Breakdown160| Category | Status | Score |161|----------|--------|-------|162| Crawlability | pass/warn/fail | XX/100 |163| Indexability | pass/warn/fail | XX/100 |164| Security | pass/warn/fail | XX/100 |165| URL Structure | pass/warn/fail | XX/100 |166| Mobile | pass/warn/fail | XX/100 |167| Core Web Vitals | pass/warn/fail | XX/100 |168| Structured Data | pass/warn/fail | XX/100 |169| JS Rendering | pass/warn/fail | XX/100 |170| IndexNow | pass/warn/fail | XX/100 |171172### Critical Issues (fix immediately)173### High Priority (fix within 1 week)174### Medium Priority (fix within 1 month)175### Low Priority (backlog)176177## DataForSEO Integration (Optional)178179If DataForSEO MCP tools are available, use `on_page_instant_pages` for real page analysis (status codes, page timing, broken links, on-page checks), `on_page_lighthouse` for Lighthouse audits (performance, accessibility, SEO scores), and `domain_analytics_technologies_domain_technologies` for technology stack detection.180181## Google API Integration (Optional)182183If Google API credentials are configured, use `python3 scripts/pagespeed_check.py <url> --json` for real PSI + CrUX field data (replaces lab-only CWV estimates), `python3 scripts/crux_history.py <url> --json` for 25-week CWV trends, and `python3 scripts/gsc_inspect.py <url> --json` for real indexation status per URL.184185## Error Handling186187| Scenario | Action |188|----------|--------|189| URL unreachable | Report connection error with status code. Suggest verifying URL, checking DNS resolution, and confirming the site is publicly accessible. |190| robots.txt not found | Note that no robots.txt was detected at the root domain. Recommend creating one with appropriate directives. Continue audit on remaining categories. |191| HTTPS not configured | Flag as a critical issue. Report whether HTTP is served without redirect, mixed content exists, or SSL certificate is missing/expired. |192| Core Web Vitals data unavailable | Note that CrUX data is not available (common for low-traffic sites). Suggest using Lighthouse lab data as a proxy and recommend increasing traffic before re-testing. |