GEO Audit Orchestration Skill
Purpose
This skill performs a comprehensive Generative Engine Optimization (GEO) audit of any website. GEO is the practice of optimizing web content so that AI systems (ChatGPT, Claude, Perplexity, Gemini, etc.) can discover, understand, cite, and recommend it. This audit measures how well a site performs across all GEO dimensions and produces an actionable improvement plan.
Key Insight
Traditional SEO optimizes for search engine rankings. GEO optimizes for AI citation and recommendation. Sites that score high on GEO metrics see 30-115% more visibility in AI-generated responses (Georgia Tech / Princeton / IIT Delhi 2024 study). The two disciplines overlap but have distinct requirements.
Audit Workflow
Phase 1: Discovery and Reconnaissance
Step 1: Fetch Homepage and Detect Business Type
Use WebFetch to retrieve the homepage at the provided URL.
Extract the following signals:
- Page title, meta description, H1 heading
- Navigation menu items (reveals site structure)
- Footer content (reveals business info, location, legal pages)
- Schema.org markup on homepage (Organization, LocalBusiness, etc.)
- Pricing page link (SaaS indicator)
- Product listing patterns (E-commerce indicator)
- Blog/resource section (Publisher indicator)
- Service pages (Agency indicator)
- Address/phone/Google Maps embed (Local business indicator)
Classify the business type using these patterns:
| Business Type |
Detection Signals |
| SaaS |
Pricing page, "Sign up" / "Free trial" CTAs, app.domain.com subdomain, feature comparison tables, integration pages |
| Local Business |
Physical address on homepage, Google Maps embed, "Near me" content, LocalBusiness schema, service area pages |
| E-commerce |
Product listings, shopping cart, product schema, category pages, price displays, "Add to cart" buttons |
| Publisher |
Blog-heavy navigation, article schema, author pages, date-based archives, RSS feeds, high content volume |
| Agency/Services |
Case studies, portfolio, "Our Work" section, team page, client logos, service descriptions |
| Hybrid |
Combination of above signals -- classify by dominant pattern |
Step 2: Crawl Sitemap and Internal Links
- Attempt to fetch
/sitemap.xml and /sitemap_index.xml.
- If sitemap exists, extract up to 50 unique page URLs prioritized by:
- Homepage (always include)
- Top-level navigation pages
- High-value pages (pricing, about, contact, key service/product pages)
- Blog posts (sample 5-10 most recent)
- Category/landing pages
- If no sitemap exists, crawl internal links from the homepage:
- Extract all
<a href> links pointing to the same domain
- Follow up to 2 levels deep
- Prioritize pages linked from main navigation
- Respect
robots.txt directives -- do not fetch disallowed paths.
- Enforce a maximum of 50 pages and a 30-second timeout per fetch.
Step 3: Collect Page-Level Data
For each page in the crawl set, record:
- URL, title, meta description, canonical URL
- H1-H6 heading structure
- Word count of main content
- Schema.org types present
- Internal/external link counts
- Images with/without alt text
- Open Graph and Twitter Card meta tags
- Response status code
- Whether the page has structured data
Phase 2: Parallel Subagent Delegation
Delegate analysis to 5 specialized subagents. Each subagent operates on the collected page data and produces a category score (0-100) plus findings.
Subagent 1: AI Citability Analysis (geo-citability)
- Analyze content blocks for quotability by AI systems
- Score passage self-containment, answer block quality, statistical density
- Identify high-value pages that could be reformatted for better AI citation
Subagent 2: Platform & Brand Analysis (geo-brand-mentions)
- Check brand presence across YouTube, Reddit, Wikipedia, LinkedIn
- Assess third-party mention volume and sentiment
- Score brand authority signals that AI models use for entity recognition
Subagent 3: Technical GEO Infrastructure (geo-crawlers + geo-llmstxt)
- Analyze robots.txt for AI crawler access
- Check for llms.txt presence and quality
- Verify meta tags, headers, and technical accessibility for AI systems
- Check page speed and rendering (JS-heavy sites are harder for AI crawlers)
Subagent 4: Content E-E-A-T Quality (geo-content)
- Evaluate Experience, Expertise, Authoritativeness, Trustworthiness signals
- Check author bios, credentials, source citations
- Assess content freshness, depth, and originality
- Verify "About" page quality and team credentials
Subagent 5: Schema & Structured Data (geo-schema)
- Validate all schema.org markup
- Check for GEO-critical schema types (FAQ, HowTo, Organization, Product, Article)
- Assess schema completeness and accuracy
- Identify missing schema opportunities
Phase 3: Score Aggregation and Report Generation
Composite GEO Score Calculation
The overall GEO Score (0-100) is a weighted average of six category scores:
| Category |
Weight |
What It Measures |
| AI Citability |
25% |
How quotable/extractable content is for AI systems |
| Brand Authority |
20% |
Third-party mentions, entity recognition signals |
| Content E-E-A-T |
20% |
Experience, Expertise, Authoritativeness, Trustworthiness |
| Technical GEO |
15% |
AI crawler access, llms.txt, rendering, speed |
| Schema & Structured Data |
10% |
Schema.org markup quality and completeness |
| Platform Optimization |
10% |
Presence on platforms AI models train on and cite |
Formula:
GEO_Score = (Citability * 0.25) + (Brand * 0.20) + (EEAT * 0.20) + (Technical * 0.15) + (Schema * 0.10) + (Platform * 0.10)
Score Interpretation
| Score Range |
Rating |
Interpretation |
| 90-100 |
Excellent |
Top-tier GEO optimization; site is highly likely to be cited by AI |
| 75-89 |
Good |
Strong GEO foundation with room for improvement |
| 60-74 |
Fair |
Moderate GEO presence; significant optimization opportunities exist |
| 40-59 |
Poor |
Weak GEO signals; AI systems may struggle to cite or recommend |
| 0-39 |
Critical |
Minimal GEO optimization; site is largely invisible to AI systems |
Issue Severity Classification
Every issue found during the audit is classified by severity:
Critical (Fix Immediately)
- All AI crawlers blocked in robots.txt
- No indexable content (JavaScript-rendered only with no SSR)
- Domain-level noindex directive
- Site returns 5xx errors on key pages
- Complete absence of any structured data
- Brand not recognized as an entity by any AI system
High (Fix Within 1 Week)
- Key AI crawlers (GPTBot, ClaudeBot, PerplexityBot) blocked
- No llms.txt file present
- Zero question-answering content blocks on key pages
- Missing Organization or LocalBusiness schema
- No author attribution on content pages
- All content behind login/paywall with no preview
Medium (Fix Within 1 Month)
- Partial AI crawler blocking (some allowed, some blocked)
- llms.txt exists but is incomplete or malformed
- Content blocks average under 50 citability score
- Missing FAQ schema on pages with FAQ content
- Thin author bios without credentials
- No Wikipedia or Reddit brand presence
Low (Optimize When Possible)
- Minor schema validation errors
- Some images missing alt text
- Content freshness issues on non-critical pages
- Missing Open Graph tags
- Suboptimal heading hierarchy on some pages
- LinkedIn company page exists but is incomplete
Output Format
Generate a file called GEO-AUDIT-REPORT.md with the following structure:
# GEO Audit Report: [Site Name]
**Audit Date:** [Date]
**URL:** [URL]
**Business Type:** [Detected Type]
**Pages Analyzed:** [Count]
---
## Executive Summary
**Overall GEO Score: [X]/100 ([Rating])**
[2-3 sentence summary of the site's GEO health, biggest strengths, and most critical gaps.]
### Score Breakdown
| Category | Score | Weight | Weighted Score |
|---|---|---|---|
| AI Citability | [X]/100 | 25% | [X] |
| Brand Authority | [X]/100 | 20% | [X] |
| Content E-E-A-T | [X]/100 | 20% | [X] |
| Technical GEO | [X]/100 | 15% | [X] |
| Schema & Structured Data | [X]/100 | 10% | [X] |
| Platform Optimization | [X]/100 | 10% | [X] |
| **Overall GEO Score** | | | **[X]/100** |
---
## Critical Issues (Fix Immediately)
[List each critical issue with specific page URLs and recommended fix]
## High Priority Issues
[List each high-priority issue with details]
## Medium Priority Issues
[List each medium-priority issue]
## Low Priority Issues
[List each low-priority issue]
---
## Category Deep Dives
### AI Citability ([X]/100)
[Detailed findings, examples of good/bad passages, rewrite suggestions]
### Brand Authority ([X]/100)
[Platform presence map, mention volume, sentiment]
### Content E-E-A-T ([X]/100)
[Author quality, source citations, freshness, depth]
### Technical GEO ([X]/100)
[Crawler access, llms.txt, rendering, headers]
### Schema & Structured Data ([X]/100)
[Schema types found, validation results, missing opportunities]
### Platform Optimization ([X]/100)
[Presence on YouTube, Reddit, Wikipedia, etc.]
---
## Quick Wins (Implement This Week)
1. [Specific, actionable quick win with expected impact]
2. [Another quick win]
3. [Another quick win]
4. [Another quick win]
5. [Another quick win]
## 30-Day Action Plan
### Week 1: [Theme]
- [ ] Action item 1
- [ ] Action item 2
### Week 2: [Theme]
- [ ] Action item 1
- [ ] Action item 2
### Week 3: [Theme]
- [ ] Action item 1
- [ ] Action item 2
### Week 4: [Theme]
- [ ] Action item 1
- [ ] Action item 2
---
## Appendix: Pages Analyzed
| URL | Title | GEO Issues |
|---|---|---|
| [url] | [title] | [issue count] |
Quality Gates
- Page Limit: Never crawl more than 50 pages per audit. Prioritize high-value pages.
- Timeout: 30-second maximum per page fetch. Skip pages that exceed this.
- Robots.txt: Always check and respect robots.txt before crawling. Note any AI-specific directives.
- Rate Limiting: Wait at least 1 second between page fetches to avoid overloading the server.
- Error Handling: Log failed fetches but continue the audit. Report fetch failures in the appendix.
- Content Type: Only analyze HTML pages. Skip PDFs, images, and other binary content.
- Deduplication: Canonicalize URLs before crawling. Skip duplicate content (e.g., HTTP vs HTTPS, www vs non-www, trailing slashes).
Business-Type-Specific Audit Adjustments
SaaS Sites
- Extra weight on: Feature comparison tables (high citability), integration pages, documentation quality
- Check for: API documentation structure, changelog pages, knowledge base organization
- Key schema: SoftwareApplication, FAQPage, HowTo
Local Businesses
- Extra weight on: NAP consistency, Google Business Profile signals, local schema
- Check for: Service area pages, location-specific content, review markup
- Key schema: LocalBusiness, GeoCoordinates, OpeningHoursSpecification
E-commerce Sites
- Extra weight on: Product descriptions (citability), comparison content, buying guides
- Check for: Product schema completeness, review aggregation, FAQ sections on product pages
- Key schema: Product, AggregateRating, Offer, BreadcrumbList
Publishers
- Extra weight on: Article quality, author credentials, source citation practices
- Check for: Article schema, author pages, publication date freshness, original research
- Key schema: Article, NewsArticle, Person (author), ClaimReview
Agency/Services
- Extra weight on: Case studies (citability), expertise demonstration, thought leadership
- Check for: Portfolio schema, team credentials, industry-specific expertise signals
- Key schema: Organization, Service, Person (team), Review
1---2name: geo-audit3description: Full website GEO+SEO audit with parallel subagent delegation. Orchestrates a comprehensive Generative Engine Optimization audit across AI citability, platform analysis, technical infrastructure, content quality, and schema markup. Produces a composite GEO Score (0-100) with prioritized action plan.4---56# GEO Audit Orchestration Skill78## Purpose910This skill performs a comprehensive Generative Engine Optimization (GEO) audit of any website. GEO is the practice of optimizing web content so that AI systems (ChatGPT, Claude, Perplexity, Gemini, etc.) can discover, understand, cite, and recommend it. This audit measures how well a site performs across all GEO dimensions and produces an actionable improvement plan.1112## Key Insight1314Traditional SEO optimizes for search engine rankings. GEO optimizes for AI citation and recommendation. Sites that score high on GEO metrics see 30-115% more visibility in AI-generated responses (Georgia Tech / Princeton / IIT Delhi 2024 study). The two disciplines overlap but have distinct requirements.1516---1718## Audit Workflow1920### Phase 1: Discovery and Reconnaissance2122**Step 1: Fetch Homepage and Detect Business Type**23241. Use WebFetch to retrieve the homepage at the provided URL.252. Extract the following signals:26 - Page title, meta description, H1 heading27 - Navigation menu items (reveals site structure)28 - Footer content (reveals business info, location, legal pages)29 - Schema.org markup on homepage (Organization, LocalBusiness, etc.)30 - Pricing page link (SaaS indicator)31 - Product listing patterns (E-commerce indicator)32 - Blog/resource section (Publisher indicator)33 - Service pages (Agency indicator)34 - Address/phone/Google Maps embed (Local business indicator)35363. Classify the business type using these patterns:3738| Business Type | Detection Signals |39|---|---|40| **SaaS** | Pricing page, "Sign up" / "Free trial" CTAs, app.domain.com subdomain, feature comparison tables, integration pages |41| **Local Business** | Physical address on homepage, Google Maps embed, "Near me" content, LocalBusiness schema, service area pages |42| **E-commerce** | Product listings, shopping cart, product schema, category pages, price displays, "Add to cart" buttons |43| **Publisher** | Blog-heavy navigation, article schema, author pages, date-based archives, RSS feeds, high content volume |44| **Agency/Services** | Case studies, portfolio, "Our Work" section, team page, client logos, service descriptions |45| **Hybrid** | Combination of above signals -- classify by dominant pattern |4647**Step 2: Crawl Sitemap and Internal Links**48491. Attempt to fetch `/sitemap.xml` and `/sitemap_index.xml`.502. If sitemap exists, extract up to 50 unique page URLs prioritized by:51 - Homepage (always include)52 - Top-level navigation pages53 - High-value pages (pricing, about, contact, key service/product pages)54 - Blog posts (sample 5-10 most recent)55 - Category/landing pages563. If no sitemap exists, crawl internal links from the homepage:57 - Extract all `<a href>` links pointing to the same domain58 - Follow up to 2 levels deep59 - Prioritize pages linked from main navigation604. Respect `robots.txt` directives -- do not fetch disallowed paths.615. Enforce a maximum of 50 pages and a 30-second timeout per fetch.6263**Step 3: Collect Page-Level Data**6465For each page in the crawl set, record:66- URL, title, meta description, canonical URL67- H1-H6 heading structure68- Word count of main content69- Schema.org types present70- Internal/external link counts71- Images with/without alt text72- Open Graph and Twitter Card meta tags73- Response status code74- Whether the page has structured data7576---7778### Phase 2: Parallel Subagent Delegation7980Delegate analysis to 5 specialized subagents. Each subagent operates on the collected page data and produces a category score (0-100) plus findings.8182**Subagent 1: AI Citability Analysis (geo-citability)**83- Analyze content blocks for quotability by AI systems84- Score passage self-containment, answer block quality, statistical density85- Identify high-value pages that could be reformatted for better AI citation8687**Subagent 2: Platform & Brand Analysis (geo-brand-mentions)**88- Check brand presence across YouTube, Reddit, Wikipedia, LinkedIn89- Assess third-party mention volume and sentiment90- Score brand authority signals that AI models use for entity recognition9192**Subagent 3: Technical GEO Infrastructure (geo-crawlers + geo-llmstxt)**93- Analyze robots.txt for AI crawler access94- Check for llms.txt presence and quality95- Verify meta tags, headers, and technical accessibility for AI systems96- Check page speed and rendering (JS-heavy sites are harder for AI crawlers)9798**Subagent 4: Content E-E-A-T Quality (geo-content)**99- Evaluate Experience, Expertise, Authoritativeness, Trustworthiness signals100- Check author bios, credentials, source citations101- Assess content freshness, depth, and originality102- Verify "About" page quality and team credentials103104**Subagent 5: Schema & Structured Data (geo-schema)**105- Validate all schema.org markup106- Check for GEO-critical schema types (FAQ, HowTo, Organization, Product, Article)107- Assess schema completeness and accuracy108- Identify missing schema opportunities109110---111112### Phase 3: Score Aggregation and Report Generation113114#### Composite GEO Score Calculation115116The overall GEO Score (0-100) is a weighted average of six category scores:117118| Category | Weight | What It Measures |119|---|---|---|120| **AI Citability** | 25% | How quotable/extractable content is for AI systems |121| **Brand Authority** | 20% | Third-party mentions, entity recognition signals |122| **Content E-E-A-T** | 20% | Experience, Expertise, Authoritativeness, Trustworthiness |123| **Technical GEO** | 15% | AI crawler access, llms.txt, rendering, speed |124| **Schema & Structured Data** | 10% | Schema.org markup quality and completeness |125| **Platform Optimization** | 10% | Presence on platforms AI models train on and cite |126127**Formula:**128```129GEO_Score = (Citability * 0.25) + (Brand * 0.20) + (EEAT * 0.20) + (Technical * 0.15) + (Schema * 0.10) + (Platform * 0.10)130```131132#### Score Interpretation133134| Score Range | Rating | Interpretation |135|---|---|---|136| 90-100 | Excellent | Top-tier GEO optimization; site is highly likely to be cited by AI |137| 75-89 | Good | Strong GEO foundation with room for improvement |138| 60-74 | Fair | Moderate GEO presence; significant optimization opportunities exist |139| 40-59 | Poor | Weak GEO signals; AI systems may struggle to cite or recommend |140| 0-39 | Critical | Minimal GEO optimization; site is largely invisible to AI systems |141142---143144## Issue Severity Classification145146Every issue found during the audit is classified by severity:147148### Critical (Fix Immediately)149- All AI crawlers blocked in robots.txt150- No indexable content (JavaScript-rendered only with no SSR)151- Domain-level noindex directive152- Site returns 5xx errors on key pages153- Complete absence of any structured data154- Brand not recognized as an entity by any AI system155156### High (Fix Within 1 Week)157- Key AI crawlers (GPTBot, ClaudeBot, PerplexityBot) blocked158- No llms.txt file present159- Zero question-answering content blocks on key pages160- Missing Organization or LocalBusiness schema161- No author attribution on content pages162- All content behind login/paywall with no preview163164### Medium (Fix Within 1 Month)165- Partial AI crawler blocking (some allowed, some blocked)166- llms.txt exists but is incomplete or malformed167- Content blocks average under 50 citability score168- Missing FAQ schema on pages with FAQ content169- Thin author bios without credentials170- No Wikipedia or Reddit brand presence171172### Low (Optimize When Possible)173- Minor schema validation errors174- Some images missing alt text175- Content freshness issues on non-critical pages176- Missing Open Graph tags177- Suboptimal heading hierarchy on some pages178- LinkedIn company page exists but is incomplete179180---181182## Output Format183184Generate a file called `GEO-AUDIT-REPORT.md` with the following structure:185186```markdown187# GEO Audit Report: [Site Name]188189**Audit Date:** [Date]190**URL:** [URL]191**Business Type:** [Detected Type]192**Pages Analyzed:** [Count]193194---195196## Executive Summary197198**Overall GEO Score: [X]/100 ([Rating])**199200[2-3 sentence summary of the site's GEO health, biggest strengths, and most critical gaps.]201202### Score Breakdown203204| Category | Score | Weight | Weighted Score |205|---|---|---|---|206| AI Citability | [X]/100 | 25% | [X] |207| Brand Authority | [X]/100 | 20% | [X] |208| Content E-E-A-T | [X]/100 | 20% | [X] |209| Technical GEO | [X]/100 | 15% | [X] |210| Schema & Structured Data | [X]/100 | 10% | [X] |211| Platform Optimization | [X]/100 | 10% | [X] |212| **Overall GEO Score** | | | **[X]/100** |213214---215216## Critical Issues (Fix Immediately)217218[List each critical issue with specific page URLs and recommended fix]219220## High Priority Issues221222[List each high-priority issue with details]223224## Medium Priority Issues225226[List each medium-priority issue]227228## Low Priority Issues229230[List each low-priority issue]231232---233234## Category Deep Dives235236### AI Citability ([X]/100)237[Detailed findings, examples of good/bad passages, rewrite suggestions]238239### Brand Authority ([X]/100)240[Platform presence map, mention volume, sentiment]241242### Content E-E-A-T ([X]/100)243[Author quality, source citations, freshness, depth]244245### Technical GEO ([X]/100)246[Crawler access, llms.txt, rendering, headers]247248### Schema & Structured Data ([X]/100)249[Schema types found, validation results, missing opportunities]250251### Platform Optimization ([X]/100)252[Presence on YouTube, Reddit, Wikipedia, etc.]253254---255256## Quick Wins (Implement This Week)2572581. [Specific, actionable quick win with expected impact]2592. [Another quick win]2603. [Another quick win]2614. [Another quick win]2625. [Another quick win]263264## 30-Day Action Plan265266### Week 1: [Theme]267- [ ] Action item 1268- [ ] Action item 2269270### Week 2: [Theme]271- [ ] Action item 1272- [ ] Action item 2273274### Week 3: [Theme]275- [ ] Action item 1276- [ ] Action item 2277278### Week 4: [Theme]279- [ ] Action item 1280- [ ] Action item 2281282---283284## Appendix: Pages Analyzed285286| URL | Title | GEO Issues |287|---|---|---|288| [url] | [title] | [issue count] |289```290291---292293## Quality Gates294295- **Page Limit:** Never crawl more than 50 pages per audit. Prioritize high-value pages.296- **Timeout:** 30-second maximum per page fetch. Skip pages that exceed this.297- **Robots.txt:** Always check and respect robots.txt before crawling. Note any AI-specific directives.298- **Rate Limiting:** Wait at least 1 second between page fetches to avoid overloading the server.299- **Error Handling:** Log failed fetches but continue the audit. Report fetch failures in the appendix.300- **Content Type:** Only analyze HTML pages. Skip PDFs, images, and other binary content.301- **Deduplication:** Canonicalize URLs before crawling. Skip duplicate content (e.g., HTTP vs HTTPS, www vs non-www, trailing slashes).302303---304305## Business-Type-Specific Audit Adjustments306307### SaaS Sites308- Extra weight on: Feature comparison tables (high citability), integration pages, documentation quality309- Check for: API documentation structure, changelog pages, knowledge base organization310- Key schema: SoftwareApplication, FAQPage, HowTo311312### Local Businesses313- Extra weight on: NAP consistency, Google Business Profile signals, local schema314- Check for: Service area pages, location-specific content, review markup315- Key schema: LocalBusiness, GeoCoordinates, OpeningHoursSpecification316317### E-commerce Sites318- Extra weight on: Product descriptions (citability), comparison content, buying guides319- Check for: Product schema completeness, review aggregation, FAQ sections on product pages320- Key schema: Product, AggregateRating, Offer, BreadcrumbList321322### Publishers323- Extra weight on: Article quality, author credentials, source citation practices324- Check for: Article schema, author pages, publication date freshness, original research325- Key schema: Article, NewsArticle, Person (author), ClaimReview326327### Agency/Services328- Extra weight on: Case studies (citability), expertise demonstration, thought leadership329- Check for: Portfolio schema, team credentials, industry-specific expertise signals330- Key schema: Organization, Service, Person (team), Review