GEO citation audit
TL;DR
|
|
| Produces |
Three sub-scores (0 to 100) plus a weighted composite, and a ranked action list |
| Default weights |
Foundations 0.30 · Answer Engine 0.20 · Generative Citation 0.50 |
| Hard gate |
Foundations below 50 stops the audit. Fix that first |
| Minimum viable run |
10 queries, 4 assistants, 20 pages sampled. Roughly two hours, no paid tooling |
| Hands off to |
sl-citation-content for content gaps, sl-entity-infrastructure for identity gaps |
Before starting
- Load
CONTEXT.md and run sl-search-surfaces if it has not run.
- Collect the inputs below. Do not start without the query list; it determines everything.
| Input |
Required |
Notes |
| Canonical entity name and aliases |
Yes |
Exactly how the brand should be named |
| Target query list |
Yes |
10 to 30, informational and commercial-investigation intent |
| Competitor list |
Yes |
3 to 10 named |
| Key page URLs |
Yes |
20 for the content sample |
| YMYL status |
Yes |
Health, finance or legal raises the authorship bar |
| Measurement tooling |
No |
Manual fallback given at every step |
- Confirm the weight profile before scoring. Changing it afterwards to flatter the result is
the fastest way to make an audit worthless.
Core framework: three dimensions
A. Foundations (0 to 100) · default weight 0.30
Can the site be crawled, indexed and understood?
| Sub |
Item |
Points |
| A1 |
Crawl and indexability, in Google and in Bing |
20 |
| A2 |
Core Web Vitals on mobile |
15 |
| A3 |
On-page hygiene: titles, headings, canonicals |
15 |
| A4 |
Schema baseline: Organization, LocalBusiness, Article |
15 |
| A5 |
Internal linking and architecture |
15 |
| A6 |
International and RTL handling where applicable |
10 |
| A7 |
Backlink health |
10 |
A1 is where most sites quietly fail for GEO. ChatGPT search runs on Bing. Teams check Google
Search Console and never open Bing Webmaster Tools, so a Bing indexing problem stays invisible while
a whole assistant cannot see them.
B. Answer Engine (0 to 100) · default weight 0.20
| Sub |
Item |
Points |
| B1 |
Snippet ownership on target queries |
30 |
| B2 |
Content structured for extraction: question H2s, direct answer paragraphs, lists, tables |
25 |
| B3 |
AEO schema coverage: FAQPage, HowTo, QAPage |
20 |
| B4 |
People Also Ask coverage |
15 |
| B5 |
Voice readiness |
10 |
C. Generative Citation (0 to 100) · default weight 0.50
| Sub |
Item |
Points |
| C1 |
LLM mention share against competitors, across assistants |
30 |
| C2 |
Entity infrastructure: Wikidata, Wikipedia, sameAs, knowledge panel |
25 |
| C3 |
Citation-worthy content, chunks an LLM can lift |
20 |
| C4 |
E-E-A-T signals: bylines, credentials, Person schema, external recognition |
15 |
| C5 |
Freshness: dateModified discipline, visible review dates, refresh cadence |
10 |
Composite
Overall = (A x W_A) + (B x W_B) + (C x W_C) where W_A + W_B + W_C = 1.0
| Profile |
W_A |
W_B |
W_C |
Use when |
| Future-weighted (default) |
0.30 |
0.20 |
0.50 |
Mature site, AI-Overview-heavy queries |
| Balanced |
0.35 |
0.30 |
0.35 |
Surface priority genuinely unclear |
| Classic |
0.50 |
0.30 |
0.20 |
Broken foundations, defer GEO |
Automatic override: Foundations below 50 forces the Classic profile regardless of the default,
and the audit reports that it did so. The tool does not recommend GEO spend on a broken site.
Workflow
[Phase 1/7: Entity] Can the machine identify the subject?
- Wikidata entry, present and correctly linked
- Wikipedia article, where notability genuinely permits. Do not advise creating one otherwise
sameAs in schema pointing at every official profile plus Wikidata
- Knowledge panel on a brand-name search
- Name, address and phone consistency across the open web
Feeds C2.
[Phase 2/7: Schema] Depth beyond the baseline
Organization or LocalBusiness on the homepage
Person for named authors, with credentials, for every editorial page
- Offering-appropriate types for products or services
Article plus author on all editorial content
dateModified present, accurate and visible to a human reader, not only in markup
Feeds A4, B3, C4, C5.
[Phase 3/7: Content] Sample 20 pages for citation-worthiness
Score each page against the five-element test from sl-citation-content:
- Entity named explicitly, no pronouns
- One definitive claim, not hedged
- A number, date or proof the model can quote
- An attributed source, inline
- Standalone readability, no "as mentioned above"
Record the count of pages carrying at least one liftable chunk. Feeds C3.
[Phase 4/7: Measure] Citation share across assistants
Run the query list through each assistant. Per query record: was the entity named, at what
position, which competitors were named, and which URL was cited.
Free method: test manually in ChatGPT search, Perplexity, Claude and Google AI Overviews.
Screenshot everything. An hour for ten queries.
Tooled method: DataForSEO ai_opt_llm_ment_*, Profound, or Otterly for continuous tracking.
Aggregate to share of citation per topic. Feeds C1.
Do not aggregate assistants into a single number without saying so. Cited-domain overlap between
ChatGPT, AI Overviews and Perplexity is low, around 11% [Profound-Citations-2026], so a blended
average hides exactly the platform-specific gap worth acting on.
[Phase 5/7: AI Overview] Zero-click exposure
For each target query: does an AI Overview trigger, is the entity cited in it, and what is the
classic rank. Trigger rates run around a quarter of searches overall and roughly half in healthcare
[Conductor-AEO-GEO-2026], and AI Overviews have been measured cutting clicks by 58%
[Ahrefs-AIO-CTR-Feb2026].
The high-value cell: an AI Overview triggers, a competitor is cited, and the site ranks in the
top ten. Ranking is present, citation is not, so the gap is content or entity, not authority. These
are the cheapest wins in the whole audit.
[Phase 6/7: Accuracy] What does the AI get wrong?
Ask each assistant directly about the brand. Log every factual error.
Errors outrank optimization. An assistant confidently stating something false is more damaging
than not being mentioned, and it needs correcting at the source the model is drawing from, not on
the site's own marketing pages.
[Phase 7/7: Map] Gaps to actions
Sort every gap into one of three buckets, because the remedy differs entirely:
| Bucket |
Symptom |
Route to |
| Identity |
The machine cannot tell who this is |
sl-entity-infrastructure |
| Content |
Identity fine, nothing liftable on the page |
sl-citation-content |
| Authority |
Both fine, still not cited |
External citations, digital PR, co-citation. Slowest to move |
Rank by leverage, not effort. State effort separately.
Best practices
- Baseline before you change anything. Without a starting citation share there is no way to show
movement, and this work is slow enough that proof matters.
- Re-measure monthly, not weekly. The signal is noisy; weekly readings mostly measure sampling.
- Keep the screenshots. Assistant outputs are not reproducible. Undated evidence is not evidence.
- Report per assistant as well as blended, given how low the overlap is.
- Say which numbers are vendor-published. Much of the measurement research in this field comes
from companies selling measurement.
Anti-patterns
| Anti-pattern |
Why it fails |
| Auditing GEO on a site failing Foundations |
Cannot be cited if it cannot be crawled |
| One blended "AI visibility score" |
Hides the platform-specific gap, which is the actionable part |
| Checking Google indexing only |
ChatGPT search runs on Bing |
| Adjusting weights after seeing the score |
Makes the audit unfalsifiable |
| Recommending a Wikipedia article for a non-notable brand |
It gets deleted, and it burns credibility |
| Treating one prompt as a measurement |
Assistants vary run to run. Sample, and record the date |
| Reporting citation share with no date |
Meaningless within a quarter |
| Promising a citation-share target |
Nobody controls these mechanisms |
Output format
Deliver as a folder:
<YYYY-MM-DD>-geo-audit/
score.md three sub-scores, weights used, composite, profile and any auto-override
citation-share.csv per query, per assistant, per competitor
entity-infra.md Wikidata, schema, sameAs, knowledge panel current state
findings.md what is true today, with evidence
priorities.md ranked actions, each tagged identity / content / authority
screenshots/ dated assistant outputs
Handoff summary block
Every audit ends with this, so a downstream skill or human can pick it up cold:
AUDIT HANDOFF
Entity: <canonical name>
Date: <YYYY-MM-DD>
Profile: <Future-weighted | Balanced | Classic> (auto-override: yes/no)
Scores: A:<0-100> B:<0-100> C:<0-100> Composite:<0-100>
Citation share: <n>% across <k> queries x <m> assistants
Top gap: <identity | content | authority>
Next skill: <sl-citation-content | sl-entity-infrastructure>
Blocking: <anything that stops work>
Figures dated: <validation date of any external statistic quoted>
Questions to ask when the brief is thin
- What ten questions should name you, in the words a real buyer would use?
- Who gets named instead today, and are they genuinely comparable?
- Where do you rank in classic search for those, right now?
- Are you in the Bing index?
- Who is the named, credentialed author on your key pages?
- When did those pages last genuinely change, not just get a touched timestamp?
- Is any of this YMYL?
- Has anyone checked what the assistants currently say about you, correct or not?
References
| Key |
Resolution |
[Aggarwal-2023] |
Aggarwal et al. GEO: Generative Engine Optimization. https://arxiv.org/abs/2311.09735 |
[Conductor-AEO-GEO-2026] |
Conductor. 2026 AEO/GEO Benchmarks Report, 21.9M searches. Vendor-published |
[Ahrefs-AIO-CTR-Feb2026] |
Ahrefs. AI Overviews reduce clicks by 58%, Feb 2026. Vendor-published |
[Profound-Citations-2026] |
Profound. AI platform citation patterns, cited-domain overlap across ChatGPT, AI Overviews and Perplexity. Vendor-published |
Full registry: references.md. Scoring rationale:
research/scoring-rubric.md.
Related skills
sl-search-surfaces, the model this audit assumes
sl-citation-content, fixes content-bucket gaps
sl-entity-infrastructure, fixes identity-bucket gaps
1---2name: sl-geo-audit3description: Audit and score a site for AI citation readiness across three dimensions: Foundations, Answer Engine and Generative Citation. Use when the user asks why AI assistants never mention their brand, wants an AI visibility audit, asks how to measure GEO, wants to know why competitors get cited instead, needs a share-of-citation baseline, or asks what to fix first to appear in AI Overviews, ChatGPT, Perplexity or Claude answers.4license: MIT5---67# GEO citation audit89## TL;DR1011| | |12|---|---|13| **Produces** | Three sub-scores (0 to 100) plus a weighted composite, and a ranked action list |14| **Default weights** | Foundations 0.30 · Answer Engine 0.20 · Generative Citation 0.50 |15| **Hard gate** | Foundations below 50 stops the audit. Fix that first |16| **Minimum viable run** | 10 queries, 4 assistants, 20 pages sampled. Roughly two hours, no paid tooling |17| **Hands off to** | `sl-citation-content` for content gaps, `sl-entity-infrastructure` for identity gaps |1819## Before starting20211. **Load [`CONTEXT.md`](../../CONTEXT.md)** and run `sl-search-surfaces` if it has not run.222. **Collect the inputs below.** Do not start without the query list; it determines everything.2324| Input | Required | Notes |25|---|---|---|26| Canonical entity name and aliases | Yes | Exactly how the brand should be named |27| Target query list | Yes | 10 to 30, informational and commercial-investigation intent |28| Competitor list | Yes | 3 to 10 named |29| Key page URLs | Yes | 20 for the content sample |30| YMYL status | Yes | Health, finance or legal raises the authorship bar |31| Measurement tooling | No | Manual fallback given at every step |32333. **Confirm the weight profile** before scoring. Changing it afterwards to flatter the result is34 the fastest way to make an audit worthless.3536## Core framework: three dimensions3738### A. Foundations (0 to 100) · default weight 0.303940Can the site be crawled, indexed and understood?4142| Sub | Item | Points |43|---|---|---|44| A1 | Crawl and indexability, **in Google and in Bing** | 20 |45| A2 | Core Web Vitals on mobile | 15 |46| A3 | On-page hygiene: titles, headings, canonicals | 15 |47| A4 | Schema baseline: Organization, LocalBusiness, Article | 15 |48| A5 | Internal linking and architecture | 15 |49| A6 | International and RTL handling where applicable | 10 |50| A7 | Backlink health | 10 |5152**A1 is where most sites quietly fail for GEO.** ChatGPT search runs on Bing. Teams check Google53Search Console and never open Bing Webmaster Tools, so a Bing indexing problem stays invisible while54a whole assistant cannot see them.5556### B. Answer Engine (0 to 100) · default weight 0.205758| Sub | Item | Points |59|---|---|---|60| B1 | Snippet ownership on target queries | 30 |61| B2 | Content structured for extraction: question H2s, direct answer paragraphs, lists, tables | 25 |62| B3 | AEO schema coverage: FAQPage, HowTo, QAPage | 20 |63| B4 | People Also Ask coverage | 15 |64| B5 | Voice readiness | 10 |6566### C. Generative Citation (0 to 100) · default weight 0.506768| Sub | Item | Points |69|---|---|---|70| C1 | LLM mention share against competitors, across assistants | 30 |71| C2 | Entity infrastructure: Wikidata, Wikipedia, `sameAs`, knowledge panel | 25 |72| C3 | Citation-worthy content, chunks an LLM can lift | 20 |73| C4 | E-E-A-T signals: bylines, credentials, `Person` schema, external recognition | 15 |74| C5 | Freshness: `dateModified` discipline, visible review dates, refresh cadence | 10 |7576### Composite7778```79Overall = (A x W_A) + (B x W_B) + (C x W_C) where W_A + W_B + W_C = 1.080```8182| Profile | W_A | W_B | W_C | Use when |83|---|---|---|---|---|84| **Future-weighted** (default) | 0.30 | 0.20 | 0.50 | Mature site, AI-Overview-heavy queries |85| **Balanced** | 0.35 | 0.30 | 0.35 | Surface priority genuinely unclear |86| **Classic** | 0.50 | 0.30 | 0.20 | Broken foundations, defer GEO |8788**Automatic override:** Foundations below 50 forces the Classic profile regardless of the default,89and the audit reports that it did so. The tool does not recommend GEO spend on a broken site.9091## Workflow9293### [Phase 1/7: Entity] Can the machine identify the subject?9495- Wikidata entry, present and correctly linked96- Wikipedia article, where notability genuinely permits. **Do not advise creating one otherwise**97- `sameAs` in schema pointing at every official profile plus Wikidata98- Knowledge panel on a brand-name search99- Name, address and phone consistency across the open web100101Feeds **C2**.102103### [Phase 2/7: Schema] Depth beyond the baseline104105- `Organization` or `LocalBusiness` on the homepage106- `Person` for named authors, with credentials, for every editorial page107- Offering-appropriate types for products or services108- `Article` plus author on all editorial content109- `dateModified` present, accurate and **visible to a human reader**, not only in markup110111Feeds **A4**, **B3**, **C4**, **C5**.112113### [Phase 3/7: Content] Sample 20 pages for citation-worthiness114115Score each page against the five-element test from `sl-citation-content`:1161171. Entity named explicitly, no pronouns1182. One definitive claim, not hedged1193. A number, date or proof the model can quote1204. An attributed source, inline1215. Standalone readability, no "as mentioned above"122123Record the count of pages carrying at least one liftable chunk. Feeds **C3**.124125### [Phase 4/7: Measure] Citation share across assistants126127Run the query list through each assistant. Per query record: was the entity named, at what128position, which competitors were named, and which URL was cited.129130**Free method:** test manually in ChatGPT search, Perplexity, Claude and Google AI Overviews.131Screenshot everything. An hour for ten queries.132**Tooled method:** DataForSEO `ai_opt_llm_ment_*`, Profound, or Otterly for continuous tracking.133134Aggregate to **share of citation** per topic. Feeds **C1**.135136> Do not aggregate assistants into a single number without saying so. Cited-domain overlap between137> ChatGPT, AI Overviews and Perplexity is low, around 11% `[Profound-Citations-2026]`, so a blended138> average hides exactly the platform-specific gap worth acting on.139140### [Phase 5/7: AI Overview] Zero-click exposure141142For each target query: does an AI Overview trigger, is the entity cited in it, and what is the143classic rank. Trigger rates run around a quarter of searches overall and roughly half in healthcare144`[Conductor-AEO-GEO-2026]`, and AI Overviews have been measured cutting clicks by 58%145`[Ahrefs-AIO-CTR-Feb2026]`.146147**The high-value cell:** an AI Overview triggers, a competitor is cited, and the site ranks in the148top ten. Ranking is present, citation is not, so the gap is content or entity, not authority. These149are the cheapest wins in the whole audit.150151### [Phase 6/7: Accuracy] What does the AI get wrong?152153Ask each assistant directly about the brand. Log every factual error.154155**Errors outrank optimization.** An assistant confidently stating something false is more damaging156than not being mentioned, and it needs correcting at the source the model is drawing from, not on157the site's own marketing pages.158159### [Phase 7/7: Map] Gaps to actions160161Sort every gap into one of three buckets, because the remedy differs entirely:162163| Bucket | Symptom | Route to |164|---|---|---|165| **Identity** | The machine cannot tell who this is | `sl-entity-infrastructure` |166| **Content** | Identity fine, nothing liftable on the page | `sl-citation-content` |167| **Authority** | Both fine, still not cited | External citations, digital PR, co-citation. Slowest to move |168169Rank by leverage, not effort. State effort separately.170171## Best practices172173- **Baseline before you change anything.** Without a starting citation share there is no way to show174 movement, and this work is slow enough that proof matters.175- **Re-measure monthly, not weekly.** The signal is noisy; weekly readings mostly measure sampling.176- **Keep the screenshots.** Assistant outputs are not reproducible. Undated evidence is not evidence.177- **Report per assistant as well as blended**, given how low the overlap is.178- **Say which numbers are vendor-published.** Much of the measurement research in this field comes179 from companies selling measurement.180181## Anti-patterns182183| Anti-pattern | Why it fails |184|---|---|185| Auditing GEO on a site failing Foundations | Cannot be cited if it cannot be crawled |186| One blended "AI visibility score" | Hides the platform-specific gap, which is the actionable part |187| Checking Google indexing only | ChatGPT search runs on Bing |188| Adjusting weights after seeing the score | Makes the audit unfalsifiable |189| Recommending a Wikipedia article for a non-notable brand | It gets deleted, and it burns credibility |190| Treating one prompt as a measurement | Assistants vary run to run. Sample, and record the date |191| Reporting citation share with no date | Meaningless within a quarter |192| Promising a citation-share target | Nobody controls these mechanisms |193194## Output format195196Deliver as a folder:197198```199<YYYY-MM-DD>-geo-audit/200 score.md three sub-scores, weights used, composite, profile and any auto-override201 citation-share.csv per query, per assistant, per competitor202 entity-infra.md Wikidata, schema, sameAs, knowledge panel current state203 findings.md what is true today, with evidence204 priorities.md ranked actions, each tagged identity / content / authority205 screenshots/ dated assistant outputs206```207208### Handoff summary block209210Every audit ends with this, so a downstream skill or human can pick it up cold:211212```213AUDIT HANDOFF214Entity: <canonical name>215Date: <YYYY-MM-DD>216Profile: <Future-weighted | Balanced | Classic> (auto-override: yes/no)217Scores: A:<0-100> B:<0-100> C:<0-100> Composite:<0-100>218Citation share: <n>% across <k> queries x <m> assistants219Top gap: <identity | content | authority>220Next skill: <sl-citation-content | sl-entity-infrastructure>221Blocking: <anything that stops work>222Figures dated: <validation date of any external statistic quoted>223```224225## Questions to ask when the brief is thin2262271. What ten questions should name you, in the words a real buyer would use?2282. Who gets named instead today, and are they genuinely comparable?2293. Where do you rank in classic search for those, right now?2304. Are you in the Bing index?2315. Who is the named, credentialed author on your key pages?2326. When did those pages last genuinely change, not just get a touched timestamp?2337. Is any of this YMYL?2348. Has anyone checked what the assistants currently say about you, correct or not?235236## References237238| Key | Resolution |239|---|---|240| `[Aggarwal-2023]` | Aggarwal et al. *GEO: Generative Engine Optimization*. https://arxiv.org/abs/2311.09735 |241| `[Conductor-AEO-GEO-2026]` | Conductor. *2026 AEO/GEO Benchmarks Report*, 21.9M searches. Vendor-published |242| `[Ahrefs-AIO-CTR-Feb2026]` | Ahrefs. *AI Overviews reduce clicks by 58%*, Feb 2026. Vendor-published |243| `[Profound-Citations-2026]` | Profound. AI platform citation patterns, cited-domain overlap across ChatGPT, AI Overviews and Perplexity. Vendor-published |244245Full registry: [`references.md`](../../references.md). Scoring rationale:246[`research/scoring-rubric.md`](../../research/scoring-rubric.md).247248## Related skills249250- `sl-search-surfaces`, the model this audit assumes251- `sl-citation-content`, fixes content-bucket gaps252- `sl-entity-infrastructure`, fixes identity-bucket gaps