Semantic Topic Clustering (v1.9.0)
SERP-overlap-driven keyword clustering for content architecture. Groups keywords
by how Google actually ranks them (shared top-10 results), not by text similarity.
Designs hub-and-spoke content clusters with internal link matrices and generates
interactive cluster map visualizations.
Scripts: Located at the plugin root scripts/ directory.
Quick Reference
| Command |
What it does |
/seo cluster plan <seed-keyword> |
Full planning workflow: expand, cluster, architect, visualize |
/seo cluster plan --from strategy |
Import from existing /seo plan output |
/seo cluster execute |
Execute plan: create content via claude-blog or output briefs |
/seo cluster map |
Regenerate the interactive cluster visualization |
Planning Workflow
Step 1: Seed Keyword Expansion
Expand the seed keyword into 30-50 variants using WebSearch:
- Related searches — Search the seed, extract "related searches" and "people also search for"
- People Also Ask (PAA) — Extract all PAA questions from SERP results
- Long-tail modifiers — Append common modifiers: "best", "how to", "vs", "for beginners", "tools", "examples", "guide", "template", "mistakes", "checklist"
- Question mining — Generate who/what/when/where/why/how variants
- Intent modifiers — Add commercial modifiers: "pricing", "review", "alternative", "comparison", "free", "top"
Deduplication: Normalize variants (lowercase, strip articles), remove exact duplicates.
Target: 30-50 unique keyword variants. If under 30, run a second expansion pass
with the top PAA questions as seeds.
Step 2: SERP Overlap Clustering
This is the core differentiator. Load references/serp-overlap-methodology.md for
the full algorithm.
Process:
- Group keywords by initial intent guess (reduces pairwise comparisons)
- For each candidate pair within a group, WebSearch both keywords
- Count shared URLs in the top 10 organic results (ignore ads, featured snippets, PAA)
- Apply thresholds:
| Shared Results |
Relationship |
Action |
| 7-10 |
Same post |
Merge into single target page |
| 4-6 |
Same cluster |
Group under same spoke cluster |
| 2-3 |
Interlink |
Place in adjacent clusters, add cross-links |
| 0-1 |
Separate |
Assign to different clusters or exclude |
Optimization: With 40 keywords, full pairwise = 780 comparisons. Instead:
- Pre-group by intent (4 groups of ~10 = 4 x 45 = 180 comparisons)
- Only cross-check group boundary keywords
- Skip pairs where both are long-tail variants of the same head term (assume same cluster)
DataForSEO integration: If DataForSEO MCP is available, use serp_organic_live_advanced
instead of WebSearch for SERP data. Run python3 scripts/dataforseo_costs.py check serp_organic_live_advanced --count N
before each batch. If "status": "needs_approval", show cost estimate and ask user.
If "status": "blocked", fall back to WebSearch.
Step 3: Intent Classification
Classify each keyword into one of four intent categories:
| Intent |
Signals |
Include in Clusters? |
| Informational |
how, what, why, guide, tutorial, learn |
Yes |
| Commercial |
best, top, review, comparison, vs, alternative |
Yes |
| Transactional |
buy, price, discount, coupon, order, sign up |
Yes |
| Navigational |
brand names, specific product names, login |
No (exclude) |
Remove navigational keywords from clustering. Flag borderline cases for
manual review. Keywords can have mixed intent (e.g., "best CRM software" is
both commercial and informational) -- classify by dominant intent.
Step 4: Hub-and-Spoke Architecture
Load references/hub-spoke-architecture.md for full specifications.
Design the cluster structure:
- Select the pillar keyword — Highest volume, broadest intent, most SERP overlap with other keywords
- Group spokes into clusters — Each cluster is a subtopic area (2-5 clusters per pillar)
- Assign posts to clusters — Each cluster gets 2-4 spoke posts
- Select templates per post — Based on intent classification:
| Intent Pattern |
Template Options |
| Informational (broad) |
ultimate-guide |
| Informational (how) |
how-to |
| Informational (list) |
listicle |
| Informational (concept) |
explainer |
| Commercial (compare) |
comparison |
| Commercial (evaluate) |
review |
| Commercial (rank) |
best-of |
| Transactional |
landing-page |
Set word count targets:
- Pillar page: 2500-4000 words
- Spoke posts: 1200-1800 words
Cannibalization check — No two posts share the same primary keyword. If SERP
overlap is 7+, merge those keywords into a single post targeting both.
Step 5: Internal Link Matrix
Design the bidirectional linking structure:
| Link Type |
Direction |
Requirement |
| Spoke to pillar |
spoke -> pillar |
Mandatory (every spoke) |
| Pillar to spoke |
pillar -> spoke |
Mandatory (every spoke) |
| Spoke to spoke (within cluster) |
spoke <-> spoke |
2-3 links per post |
| Cross-cluster |
spoke -> spoke (other cluster) |
0-1 links per post |
Rules:
- Every post must have minimum 3 incoming internal links
- No orphan pages (every post reachable from pillar in 2 clicks)
- Anchor text must use target keyword or close variant (no "click here")
- Link placement: within body content, not just navigation/sidebar
Generate the link matrix as a JSON adjacency list:
{
"links": [
{ "from": "pillar", "to": "cluster-0-post-0", "type": "mandatory", "anchor": "keyword" },
{ "from": "cluster-0-post-0", "to": "pillar", "type": "mandatory", "anchor": "keyword" }
]
}
Step 6: Interactive Cluster Map
Generate cluster-map.html using the template at templates/cluster-map.html.
- Read the template file
- Build the
CLUSTER_DATA JSON object from the cluster plan:{
pillar: { title, keyword, volume, template, wordCount, url },
clusters: [{ name, color, posts: [{ title, keyword, volume, template, wordCount, url, status }] }],
links: [{ from, to, type }],
meta: { totalPosts, totalClusters, totalLinks, estimatedWords }
}
- Replace the
CLUSTER_DATA placeholder in the template with the actual JSON
- Write the completed HTML file to the output directory
- Inform user: "Open
cluster-map.html in a browser to explore the interactive cluster map."
Strategy Import
When invoked with --from strategy:
- Look for the most recent
/seo plan output in the current directory (search for
files matching *SEO*Plan*, *strategy*, *content-strategy*)
- Parse markdown tables for: keywords, page types, content pillars, URL structures
- Validate extracted data: check for duplicates, missing keywords, incomplete entries
- Enrich with SERP data: run SERP overlap analysis on extracted keywords
- Build cluster plan using the imported keywords as the starting set (skip Step 1)
If no strategy file is found, prompt the user: "No existing SEO plan found in the
current directory. Run /seo plan first, or provide a seed keyword for fresh clustering."
Execution Workflow
When /seo cluster execute is invoked:
Check for claude-blog
Test: Does ~/.claude/skills/blog/SKILL.md exist?
If claude-blog IS installed:
- Load
references/execution-workflow.md for the full algorithm
- Read
cluster-plan.json from the current directory
- Check for resume state: scan output directory for already-written posts
- Execute in priority order: pillar first, then spokes by volume (highest first)
- For each post, invoke the
blog-write skill with cluster context:
- Cluster role (pillar or spoke)
- Position in cluster (cluster index, post index)
- Target keyword and secondary keywords
- Template type and word count target
- Internal links to include (with anchors)
- Links to receive from future posts (placeholder markers)
- After each post is written, scan previous posts for backward link placeholders
and inject the new post's URL
- After all posts are written, generate the cluster scorecard
If claude-blog is NOT installed:
- Generate detailed content briefs for each post in the cluster plan
- Each brief includes:
- Title and meta description
- Primary keyword and secondary keywords
- Template type and suggested structure (H2/H3 outline)
- Word count target
- Internal links to include (with anchor text)
- Key points to cover
- Competing pages to differentiate from
- Write briefs to
cluster-briefs/ directory as individual markdown files
- Inform user: "Install claude-blog
to auto-create content. Briefs saved to
cluster-briefs/."
Cluster Scorecard
Post-execution quality report. Run automatically after /seo cluster execute or
on demand via analysis of the output directory.
| Metric |
Target |
How Measured |
| Coverage |
100% |
Posts written / posts planned |
| Link Density |
3+ per post |
Count internal links per post |
| Orphan Pages |
0 |
Posts with < 1 incoming link |
| Cannibalization |
0 conflicts |
Check for duplicate primary keywords |
| Image Count |
1+ per post |
Posts with at least one image |
| Pillar Links |
100% |
All spokes link to pillar and vice versa |
| Cross-Links |
80%+ |
Recommended spoke-to-spoke links implemented |
| Content Gaps |
0 |
Planned posts that were skipped or incomplete |
Map Regeneration
When /seo cluster map is invoked:
- Read
cluster-plan.json from the current directory
- Scan output directory and update post statuses (planned vs written)
- Regenerate
cluster-map.html with updated statuses
- Report: posts written vs planned, link completion percentage
Output Files
All outputs are written to the current working directory:
| File |
Description |
cluster-plan.json |
Machine-readable cluster plan (full data) |
cluster-plan.md |
Human-readable cluster plan summary |
cluster-map.html |
Interactive SVG visualization |
cluster-briefs/ |
Content briefs (if no claude-blog) |
cluster-scorecard.md |
Post-execution quality report |
Cross-Skill Integration
| Skill |
Relationship |
seo-plan |
Import source: strategy import reads seo-plan output |
seo-content |
Quality check: E-E-A-T validation of generated content |
seo-schema |
Schema markup: Article, BreadcrumbList, ItemList for cluster pages |
seo-dataforseo |
Data source: SERP data when DataForSEO MCP is available |
seo-google |
Reporting: generate PDF report of cluster plan and scorecard |
After cluster planning or execution completes, offer:
"Generate a PDF report? Use /seo google report"
Error Handling
| Error |
Cause |
Resolution |
| "No seed keyword provided" |
Missing argument |
Prompt user for seed keyword or URL |
| "Insufficient keyword variants" |
Expansion yielded < 15 keywords |
Run second expansion pass with PAA questions |
| "SERP data unavailable" |
WebSearch and DataForSEO both failing |
Retry after 30s; if persistent, use intent-only clustering with warning |
| "No strategy file found" |
--from strategy but no plan exists |
Prompt user to run /seo plan first |
| "cluster-plan.json not found" |
Execute without planning |
Prompt user to run /seo cluster plan first |
| "claude-blog not installed" |
Execute attempted without blog skill |
Generate content briefs instead; suggest installation |
| "DataForSEO budget exceeded" |
Cost check returned "blocked" |
Fall back to WebSearch; inform user |
| "Duplicate primary keywords" |
Cannibalization detected |
Merge affected posts or reassign keywords |
| "Orphan page detected" |
Post missing incoming links |
Add links from nearest cluster siblings |
| "Resume state corrupted" |
Mismatch between plan and output |
Rebuild state from output directory scan |
Security
- All URLs fetched via
python3 scripts/render_page.py --mode auto (SPA-aware SSRF protection via url_safety)
- No credentials stored or transmitted
- Output files contain no PII or API keys
- DataForSEO cost checks run before every API call
FLOW Framework Integration
For prompt-guided keyword research and gap analysis, use /seo flow find [url|topic] — FLOW's 5 find-stage prompts complement the SERP-overlap clustering methodology with structured discovery prompts.
Source: hashgraph-online/awesome-codex-plugins → plugins/avalonreset/seo-dungeon/skills/seo-cluster/SKILL.md
1---2name: seo-cluster3description: > SERP-based semantic topic clustering for content architecture planning. Groups keywords by actual Google SERP overlap (not text similarity), designs hub-and-spoke content clusters with internal link matrices, and generates interactive visualizations. Optionally executes content creation if claude-blog is installed. Use when user says "topic cluster", "content cluster", "semantic clustering", "pillar page", "hub and spoke", "content architecture", "keyword grouping", or "cluster plan".4---5
6
7# Semantic Topic Clustering (v1.9.0)
8
9SERP-overlap-driven keyword clustering for content architecture. Groups keywords
10by how Google actually ranks them (shared top-10 results), not by text similarity.
11Designs hub-and-spoke content clusters with internal link matrices and generates
12interactive cluster map visualizations.
13
14**Scripts:** Located at the plugin root `scripts/` directory.
15
16---
17
18## Quick Reference
19
20| Command | What it does |
21|---------|-------------|
22| `/seo cluster plan <seed-keyword>` | Full planning workflow: expand, cluster, architect, visualize |
23| `/seo cluster plan --from strategy` | Import from existing `/seo plan` output |
24| `/seo cluster execute` | Execute plan: create content via claude-blog or output briefs |
25| `/seo cluster map` | Regenerate the interactive cluster visualization |
26
27---
28
29## Planning Workflow
30
31### Step 1: Seed Keyword Expansion
32
33Expand the seed keyword into 30-50 variants using WebSearch:
34
351. **Related searches** — Search the seed, extract "related searches" and "people also search for"
362. **People Also Ask (PAA)** — Extract all PAA questions from SERP results
373. **Long-tail modifiers** — Append common modifiers: "best", "how to", "vs", "for beginners", "tools", "examples", "guide", "template", "mistakes", "checklist"
384. **Question mining** — Generate who/what/when/where/why/how variants
395. **Intent modifiers** — Add commercial modifiers: "pricing", "review", "alternative", "comparison", "free", "top"
40
41**Deduplication:** Normalize variants (lowercase, strip articles), remove exact duplicates.
42Target: 30-50 unique keyword variants. If under 30, run a second expansion pass
43with the top PAA questions as seeds.
44
45### Step 2: SERP Overlap Clustering
46
47This is the core differentiator. Load `references/serp-overlap-methodology.md` for
48the full algorithm.
49
50**Process:**
511. Group keywords by initial intent guess (reduces pairwise comparisons)
522. For each candidate pair within a group, WebSearch both keywords
533. Count shared URLs in the top 10 organic results (ignore ads, featured snippets, PAA)
544. Apply thresholds:
55
56| Shared Results | Relationship | Action |
57|---------------|-------------|--------|
58| 7-10 | Same post | Merge into single target page |
59| 4-6 | Same cluster | Group under same spoke cluster |
60| 2-3 | Interlink | Place in adjacent clusters, add cross-links |
61| 0-1 | Separate | Assign to different clusters or exclude |
62
63**Optimization:** With 40 keywords, full pairwise = 780 comparisons. Instead:
64- Pre-group by intent (4 groups of ~10 = 4 x 45 = 180 comparisons)
65- Only cross-check group boundary keywords
66- Skip pairs where both are long-tail variants of the same head term (assume same cluster)
67
68**DataForSEO integration:** If DataForSEO MCP is available, use `serp_organic_live_advanced`
69instead of WebSearch for SERP data. Run `python3 scripts/dataforseo_costs.py check serp_organic_live_advanced --count N`
70before each batch. If `"status": "needs_approval"`, show cost estimate and ask user.
71If `"status": "blocked"`, fall back to WebSearch.
72
73### Step 3: Intent Classification
74
75Classify each keyword into one of four intent categories:
76
77| Intent | Signals | Include in Clusters? |
78|--------|---------|---------------------|
79| Informational | how, what, why, guide, tutorial, learn | Yes |
80| Commercial | best, top, review, comparison, vs, alternative | Yes |
81| Transactional | buy, price, discount, coupon, order, sign up | Yes |
82| Navigational | brand names, specific product names, login | No (exclude) |
83
84Remove navigational keywords from clustering. Flag borderline cases for
85manual review. Keywords can have mixed intent (e.g., "best CRM software" is
86both commercial and informational) -- classify by dominant intent.
87
88### Step 4: Hub-and-Spoke Architecture
89
90Load `references/hub-spoke-architecture.md` for full specifications.
91
92**Design the cluster structure:**
93
941. **Select the pillar keyword** — Highest volume, broadest intent, most SERP overlap with other keywords
952. **Group spokes into clusters** — Each cluster is a subtopic area (2-5 clusters per pillar)
963. **Assign posts to clusters** — Each cluster gets 2-4 spoke posts
974. **Select templates per post** — Based on intent classification:
98
99| Intent Pattern | Template Options |
100|---------------|-----------------|
101| Informational (broad) | ultimate-guide |
102| Informational (how) | how-to |
103| Informational (list) | listicle |
104| Informational (concept) | explainer |
105| Commercial (compare) | comparison |
106| Commercial (evaluate) | review |
107| Commercial (rank) | best-of |
108| Transactional | landing-page |
109
1105. **Set word count targets:**
111 - Pillar page: 2500-4000 words
112 - Spoke posts: 1200-1800 words
113
1146. **Cannibalization check** — No two posts share the same primary keyword. If SERP
115 overlap is 7+, merge those keywords into a single post targeting both.
116
117### Step 5: Internal Link Matrix
118
119Design the bidirectional linking structure:
120
121| Link Type | Direction | Requirement |
122|-----------|-----------|-------------|
123| Spoke to pillar | spoke -> pillar | Mandatory (every spoke) |
124| Pillar to spoke | pillar -> spoke | Mandatory (every spoke) |
125| Spoke to spoke (within cluster) | spoke <-> spoke | 2-3 links per post |
126| Cross-cluster | spoke -> spoke (other cluster) | 0-1 links per post |
127
128**Rules:**
129- Every post must have minimum 3 incoming internal links
130- No orphan pages (every post reachable from pillar in 2 clicks)
131- Anchor text must use target keyword or close variant (no "click here")
132- Link placement: within body content, not just navigation/sidebar
133
134Generate the link matrix as a JSON adjacency list:
135```json
136{
137 "links": [
138 { "from": "pillar", "to": "cluster-0-post-0", "type": "mandatory", "anchor": "keyword" },
139 { "from": "cluster-0-post-0", "to": "pillar", "type": "mandatory", "anchor": "keyword" }
140 ]
141}
142```
143
144### Step 6: Interactive Cluster Map
145
146Generate `cluster-map.html` using the template at `templates/cluster-map.html`.
147
1481. Read the template file
1492. Build the `CLUSTER_DATA` JSON object from the cluster plan:
150 ```javascript
151 {
152 pillar: { title, keyword, volume, template, wordCount, url },
153 clusters: [{ name, color, posts: [{ title, keyword, volume, template, wordCount, url, status }] }],
154 links: [{ from, to, type }],
155 meta: { totalPosts, totalClusters, totalLinks, estimatedWords }
156 }
157 ```
1583. Replace the `CLUSTER_DATA` placeholder in the template with the actual JSON
1594. Write the completed HTML file to the output directory
1605. Inform user: "Open `cluster-map.html` in a browser to explore the interactive cluster map."
161
162---
163
164## Strategy Import
165
166When invoked with `--from strategy`:
167
1681. Look for the most recent `/seo plan` output in the current directory (search for
169 files matching `*SEO*Plan*`, `*strategy*`, `*content-strategy*`)
1702. Parse markdown tables for: keywords, page types, content pillars, URL structures
1713. Validate extracted data: check for duplicates, missing keywords, incomplete entries
1724. Enrich with SERP data: run SERP overlap analysis on extracted keywords
1735. Build cluster plan using the imported keywords as the starting set (skip Step 1)
174
175If no strategy file is found, prompt the user: "No existing SEO plan found in the
176current directory. Run `/seo plan` first, or provide a seed keyword for fresh clustering."
177
178---
179
180## Execution Workflow
181
182When `/seo cluster execute` is invoked:
183
184### Check for claude-blog
185
186```
187Test: Does ~/.claude/skills/blog/SKILL.md exist?
188```
189
190**If claude-blog IS installed:**
191
1921. Load `references/execution-workflow.md` for the full algorithm
1932. Read `cluster-plan.json` from the current directory
1943. Check for resume state: scan output directory for already-written posts
1954. Execute in priority order: pillar first, then spokes by volume (highest first)
1965. For each post, invoke the `blog-write` skill with cluster context:
197 - Cluster role (pillar or spoke)
198 - Position in cluster (cluster index, post index)
199 - Target keyword and secondary keywords
200 - Template type and word count target
201 - Internal links to include (with anchors)
202 - Links to receive from future posts (placeholder markers)
2036. After each post is written, scan previous posts for backward link placeholders
204 and inject the new post's URL
2057. After all posts are written, generate the cluster scorecard
206
207**If claude-blog is NOT installed:**
208
2091. Generate detailed content briefs for each post in the cluster plan
2102. Each brief includes:
211 - Title and meta description
212 - Primary keyword and secondary keywords
213 - Template type and suggested structure (H2/H3 outline)
214 - Word count target
215 - Internal links to include (with anchor text)
216 - Key points to cover
217 - Competing pages to differentiate from
2183. Write briefs to `cluster-briefs/` directory as individual markdown files
2194. Inform user: "Install [claude-blog](https://github.com/AgriciDaniel/claude-blog)
220 to auto-create content. Briefs saved to `cluster-briefs/`."
221
222---
223
224## Cluster Scorecard
225
226Post-execution quality report. Run automatically after `/seo cluster execute` or
227on demand via analysis of the output directory.
228
229| Metric | Target | How Measured |
230|--------|--------|-------------|
231| Coverage | 100% | Posts written / posts planned |
232| Link Density | 3+ per post | Count internal links per post |
233| Orphan Pages | 0 | Posts with < 1 incoming link |
234| Cannibalization | 0 conflicts | Check for duplicate primary keywords |
235| Image Count | 1+ per post | Posts with at least one image |
236| Pillar Links | 100% | All spokes link to pillar and vice versa |
237| Cross-Links | 80%+ | Recommended spoke-to-spoke links implemented |
238| Content Gaps | 0 | Planned posts that were skipped or incomplete |
239
240---
241
242## Map Regeneration
243
244When `/seo cluster map` is invoked:
245
2461. Read `cluster-plan.json` from the current directory
2472. Scan output directory and update post statuses (planned vs written)
2483. Regenerate `cluster-map.html` with updated statuses
2494. Report: posts written vs planned, link completion percentage
250
251---
252
253## Output Files
254
255All outputs are written to the current working directory:
256
257| File | Description |
258|------|-------------|
259| `cluster-plan.json` | Machine-readable cluster plan (full data) |
260| `cluster-plan.md` | Human-readable cluster plan summary |
261| `cluster-map.html` | Interactive SVG visualization |
262| `cluster-briefs/` | Content briefs (if no claude-blog) |
263| `cluster-scorecard.md` | Post-execution quality report |
264
265---
266
267## Cross-Skill Integration
268
269| Skill | Relationship |
270|-------|-------------|
271| `seo-plan` | Import source: strategy import reads seo-plan output |
272| `seo-content` | Quality check: E-E-A-T validation of generated content |
273| `seo-schema` | Schema markup: Article, BreadcrumbList, ItemList for cluster pages |
274| `seo-dataforseo` | Data source: SERP data when DataForSEO MCP is available |
275| `seo-google` | Reporting: generate PDF report of cluster plan and scorecard |
276
277After cluster planning or execution completes, offer:
278"Generate a PDF report? Use `/seo google report`"
279
280---
281
282## Error Handling
283
284| Error | Cause | Resolution |
285|-------|-------|------------|
286| "No seed keyword provided" | Missing argument | Prompt user for seed keyword or URL |
287| "Insufficient keyword variants" | Expansion yielded < 15 keywords | Run second expansion pass with PAA questions |
288| "SERP data unavailable" | WebSearch and DataForSEO both failing | Retry after 30s; if persistent, use intent-only clustering with warning |
289| "No strategy file found" | `--from strategy` but no plan exists | Prompt user to run `/seo plan` first |
290| "cluster-plan.json not found" | Execute without planning | Prompt user to run `/seo cluster plan` first |
291| "claude-blog not installed" | Execute attempted without blog skill | Generate content briefs instead; suggest installation |
292| "DataForSEO budget exceeded" | Cost check returned "blocked" | Fall back to WebSearch; inform user |
293| "Duplicate primary keywords" | Cannibalization detected | Merge affected posts or reassign keywords |
294| "Orphan page detected" | Post missing incoming links | Add links from nearest cluster siblings |
295| "Resume state corrupted" | Mismatch between plan and output | Rebuild state from output directory scan |
296
297---
298
299## Security
300
301- All URLs fetched via `python3 scripts/render_page.py --mode auto` (SPA-aware SSRF protection via `url_safety`)
302- No credentials stored or transmitted
303- Output files contain no PII or API keys
304- DataForSEO cost checks run before every API call
305
306## FLOW Framework Integration
307
308For prompt-guided keyword research and gap analysis, use `/seo flow find [url|topic]` — FLOW's 5 find-stage prompts complement the SERP-overlap clustering methodology with structured discovery prompts.
309
310---
311
312**Source:** [`hashgraph-online/awesome-codex-plugins`](https://github.com/hashgraph-online/awesome-codex-plugins) → `plugins/avalonreset/seo-dungeon/skills/seo-cluster/SKILL.md`