Site Content Catalog
Crawl a website's sitemap and blog to build a complete content inventory — every page cataloged with URL, title, date, content type, and topic cluster. Groups content by category, identifies publishing patterns, and optionally deep-analyzes top pages.
Quick Start
# Basic content inventory
python3 scripts/catalog_content.py --domain "example.com"
# With deep analysis of top 20 pages
python3 scripts/catalog_content.py --domain "example.com" --deep-analyze 20
# Output to specific file
python3 scripts/catalog_content.py --domain "example.com" --output content-inventory.json
Inputs
| Parameter |
Required |
Default |
Description |
| domain |
Yes |
— |
Domain to catalog (e.g., "example.com") |
| deep-analyze |
No |
0 |
Number of top pages to deep-read for content analysis |
| output |
No |
stdout |
Path to save JSON output |
| include-non-blog |
No |
true |
Also catalog landing pages, docs, etc. (not just blog) |
Cost
- Sitemap/RSS crawling: Free (direct HTTP requests)
- Apify sitemap extractor (fallback): ~$0.50 per site
- Deep analysis: Free (WebFetch on individual pages)
Process
Phase 1: Discover All Pages
The script attempts multiple methods to find all pages on a site, in order:
A) Sitemap.xml
- Fetch
https://[domain]/sitemap.xml
- If it's a sitemap index, recursively fetch all child sitemaps
- Common alternate locations:
/sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml
- Check
robots.txt for Sitemap: directives
B) RSS/Atom Feeds
- Check
/feed, /rss, /atom.xml, /blog/feed, etc.
- Extract posts with titles, dates, and URLs
- RSS typically only surfaces recent content (last 10-50 posts)
C) Blog Index Crawl
- Fetch
/blog, /resources, /insights, /news, /articles
- Extract links from the page
- Follow pagination if present (
/blog/page/2, ?page=2, etc.)
D) Site: Search (fallback)
- WebSearch:
site:[domain] to estimate total indexed pages
- WebSearch:
site:[domain]/blog to find blog content
- WebSearch:
site:[domain] intitle: to discover page title patterns
E) Apify Sitemap Extractor (fallback for JS-heavy sites)
- Actor:
onescales/sitemap-url-extractor
- Use when sitemap.xml is missing and the site is JS-rendered
Phase 2: Classify Each Page
For each discovered URL, classify by:
Content Type
Classify based on URL patterns and page titles:
| Type |
URL Patterns |
Examples |
blog-post |
/blog/, /posts/, /articles/ |
How-to guides, opinion pieces |
case-study |
/case-study/, /customers/, /success-stories/ |
Customer stories |
comparison |
/vs/, /compare/, /alternative/ |
X vs Y pages |
landing-page |
/solutions/, /use-cases/, /for-/ |
Product marketing pages |
docs |
/docs/, /help/, /documentation/, /api/ |
Technical documentation |
changelog |
/changelog/, /releases/, /whats-new/ |
Product updates |
pricing |
/pricing/ |
Pricing page |
about |
/about/, /team/, /careers/ |
Company pages |
legal |
/privacy/, /terms/, /security/ |
Legal/compliance |
resource |
/resources/, /guides/, /ebooks/, /webinars/ |
Gated/downloadable content |
glossary |
/glossary/, /dictionary/, /terms/ |
SEO glossary pages |
integration |
/integrations/, /apps/, /marketplace/ |
Integration pages |
other |
— |
Anything else |
Topic Cluster
Group by extracting topic signals from URL slugs and titles:
- Extract keywords from URL path segments
- Group similar keywords into clusters (e.g., "aws-cost", "cloud-spending", "finops" → "Cloud Cost Management")
- Use simple keyword co-occurrence for clustering
Phase 3: Analyze Publishing Patterns
From the dated content (primarily blog posts):
- Total content pieces by type
- Publishing frequency: Posts per month over last 12 months
- Trend: Increasing, decreasing, or stable output
- Recency: Date of most recent publish
- Author diversity: Unique authors (if extractable from RSS)
Phase 4: Deep Analysis (Optional)
If --deep-analyze N is specified, fetch the top N pages (prioritizing blog posts) and extract:
- Word count (approximate)
- Target keyword (inferred from title + H1 + URL)
- Funnel stage: TOFU (awareness), MOFU (consideration), BOFU (decision)
- Content depth: Shallow (<500 words), Medium (500-1500), Deep (1500+)
- Has images/video: Boolean
- Has CTA: Boolean (detected by common CTA patterns)
- Internal links count
Phase 5: Output
JSON Output (default)
{
"domain": "example.com",
"crawl_date": "2026-02-25",
"total_pages": 347,
"discovery_methods": ["sitemap.xml", "rss"],
"pages": [
{
"url": "https://example.com/blog/reduce-aws-costs",
"title": "How to Reduce Your AWS Bill by 40%",
"date": "2025-11-15",
"type": "blog-post",
"topic_cluster": "Cloud Cost Optimization",
"deep_analysis": {
"word_count": 2100,
"target_keyword": "reduce aws costs",
"funnel_stage": "TOFU",
"content_depth": "deep",
"has_images": true,
"has_cta": true
}
}
],
"summary": {
"by_type": {"blog-post": 89, "landing-page": 23, "case-study": 12, ...},
"by_topic": {"Cloud Cost Optimization": 34, "FinOps": 18, ...},
"publishing_cadence": {
"posts_per_month_avg": 4.2,
"trend": "increasing",
"most_recent": "2026-02-20"
}
}
}
Markdown Summary (also generated)
# Content Inventory: example.com
**Crawled:** 2026-02-25 | **Total pages:** 347
## Content by Type
| Type | Count | % |
|------|-------|---|
| Blog Posts | 89 | 25.6% |
| Landing Pages | 23 | 6.6% |
| ...
## Content by Topic Cluster
| Topic | Posts | Most Recent |
|-------|-------|-------------|
| Cloud Cost Optimization | 34 | 2026-02-20 |
| ...
## Publishing Cadence
- Average: 4.2 posts/month
- Trend: Increasing (3.1 → 5.4 over last 6 months)
- Most recent: 2026-02-20
## Full Catalog
| # | Date | Type | Topic | Title | URL |
|---|------|------|-------|-------|-----|
| 1 | 2026-02-20 | blog-post | Cloud Cost | How to Reduce... | https://... |
Tips
- Sitemap.xml is the best source. Most well-maintained sites have one. If missing, it's itself an SEO signal (negative).
- RSS only shows recent content. If you need the full catalog, sitemap is essential. RSS is supplementary.
- Deep analysis is optional but valuable. Use it when feeding into brand-voice-extractor or when you need funnel stage mapping.
- JS-rendered sites may need the Apify fallback. Signs: sitemap.xml returns HTML, or blog page returns mostly JavaScript.
- Combine with seo-domain-analyzer to overlay traffic data on the content inventory — see which content actually performs.
Dependencies
- Python 3.8+
requests library (pip install requests)
APIFY_API_TOKEN env var (only for Apify fallback mode)
1---2name: site-content-catalog3description: Crawl a website's sitemap and blog index to build a complete content inventory. Lists every page with URL, title, publish date, content type, and topic cluster. Groups content by category and topic. Optionally deep-reads top N pages for quality analysis and funnel stage tagging. Use before SEO audits, content gap analysis, or brand voice extraction.4---5
6# Site Content Catalog
7
8Crawl a website's sitemap and blog to build a complete content inventory — every page cataloged with URL, title, date, content type, and topic cluster. Groups content by category, identifies publishing patterns, and optionally deep-analyzes top pages.
9
10## Quick Start
11
12```bash
13# Basic content inventory
14python3 scripts/catalog_content.py --domain "example.com"
15
16# With deep analysis of top 20 pages
17python3 scripts/catalog_content.py --domain "example.com" --deep-analyze 20
18
19# Output to specific file
20python3 scripts/catalog_content.py --domain "example.com" --output content-inventory.json
21```
22
23## Inputs
24
25| Parameter | Required | Default | Description |
26|-----------|----------|---------|-------------|
27| domain | Yes | — | Domain to catalog (e.g., "example.com") |
28| deep-analyze | No | 0 | Number of top pages to deep-read for content analysis |
29| output | No | stdout | Path to save JSON output |
30| include-non-blog | No | true | Also catalog landing pages, docs, etc. (not just blog) |
31
32## Cost
33
34- **Sitemap/RSS crawling:** Free (direct HTTP requests)
35- **Apify sitemap extractor (fallback):** ~$0.50 per site
36- **Deep analysis:** Free (WebFetch on individual pages)
37
38## Process
39
40### Phase 1: Discover All Pages
41
42The script attempts multiple methods to find all pages on a site, in order:
43
44#### A) Sitemap.xml
451. Fetch `https://[domain]/sitemap.xml`
462. If it's a sitemap index, recursively fetch all child sitemaps
473. Common alternate locations: `/sitemap_index.xml`, `/sitemap-index.xml`, `/wp-sitemap.xml`
484. Check `robots.txt` for `Sitemap:` directives
49
50#### B) RSS/Atom Feeds
511. Check `/feed`, `/rss`, `/atom.xml`, `/blog/feed`, etc.
522. Extract posts with titles, dates, and URLs
533. RSS typically only surfaces recent content (last 10-50 posts)
54
55#### C) Blog Index Crawl
561. Fetch `/blog`, `/resources`, `/insights`, `/news`, `/articles`
572. Extract links from the page
583. Follow pagination if present (`/blog/page/2`, `?page=2`, etc.)
59
60#### D) Site: Search (fallback)
611. WebSearch: `site:[domain]` to estimate total indexed pages
622. WebSearch: `site:[domain]/blog` to find blog content
633. WebSearch: `site:[domain] intitle:` to discover page title patterns
64
65#### E) Apify Sitemap Extractor (fallback for JS-heavy sites)
66- Actor: `onescales/sitemap-url-extractor`
67- Use when sitemap.xml is missing and the site is JS-rendered
68
69### Phase 2: Classify Each Page
70
71For each discovered URL, classify by:
72
73#### Content Type
74Classify based on URL patterns and page titles:
75
76| Type | URL Patterns | Examples |
77|------|-------------|----------|
78| `blog-post` | `/blog/`, `/posts/`, `/articles/` | How-to guides, opinion pieces |
79| `case-study` | `/case-study/`, `/customers/`, `/success-stories/` | Customer stories |
80| `comparison` | `/vs/`, `/compare/`, `/alternative/` | X vs Y pages |
81| `landing-page` | `/solutions/`, `/use-cases/`, `/for-/` | Product marketing pages |
82| `docs` | `/docs/`, `/help/`, `/documentation/`, `/api/` | Technical documentation |
83| `changelog` | `/changelog/`, `/releases/`, `/whats-new/` | Product updates |
84| `pricing` | `/pricing/` | Pricing page |
85| `about` | `/about/`, `/team/`, `/careers/` | Company pages |
86| `legal` | `/privacy/`, `/terms/`, `/security/` | Legal/compliance |
87| `resource` | `/resources/`, `/guides/`, `/ebooks/`, `/webinars/` | Gated/downloadable content |
88| `glossary` | `/glossary/`, `/dictionary/`, `/terms/` | SEO glossary pages |
89| `integration` | `/integrations/`, `/apps/`, `/marketplace/` | Integration pages |
90| `other` | — | Anything else |
91
92#### Topic Cluster
93Group by extracting topic signals from URL slugs and titles:
94- Extract keywords from URL path segments
95- Group similar keywords into clusters (e.g., "aws-cost", "cloud-spending", "finops" → "Cloud Cost Management")
96- Use simple keyword co-occurrence for clustering
97
98### Phase 3: Analyze Publishing Patterns
99
100From the dated content (primarily blog posts):
101- **Total content pieces** by type
102- **Publishing frequency:** Posts per month over last 12 months
103- **Trend:** Increasing, decreasing, or stable output
104- **Recency:** Date of most recent publish
105- **Author diversity:** Unique authors (if extractable from RSS)
106
107### Phase 4: Deep Analysis (Optional)
108
109If `--deep-analyze N` is specified, fetch the top N pages (prioritizing blog posts) and extract:
110- **Word count** (approximate)
111- **Target keyword** (inferred from title + H1 + URL)
112- **Funnel stage:** TOFU (awareness), MOFU (consideration), BOFU (decision)
113- **Content depth:** Shallow (<500 words), Medium (500-1500), Deep (1500+)
114- **Has images/video:** Boolean
115- **Has CTA:** Boolean (detected by common CTA patterns)
116- **Internal links count**
117
118### Phase 5: Output
119
120#### JSON Output (default)
121```json
122{
123 "domain": "example.com",
124 "crawl_date": "2026-02-25",
125 "total_pages": 347,
126 "discovery_methods": ["sitemap.xml", "rss"],
127 "pages": [
128 {
129 "url": "https://example.com/blog/reduce-aws-costs",
130 "title": "How to Reduce Your AWS Bill by 40%",
131 "date": "2025-11-15",
132 "type": "blog-post",
133 "topic_cluster": "Cloud Cost Optimization",
134 "deep_analysis": {
135 "word_count": 2100,
136 "target_keyword": "reduce aws costs",
137 "funnel_stage": "TOFU",
138 "content_depth": "deep",
139 "has_images": true,
140 "has_cta": true
141 }
142 }
143 ],
144 "summary": {
145 "by_type": {"blog-post": 89, "landing-page": 23, "case-study": 12, ...},
146 "by_topic": {"Cloud Cost Optimization": 34, "FinOps": 18, ...},
147 "publishing_cadence": {
148 "posts_per_month_avg": 4.2,
149 "trend": "increasing",
150 "most_recent": "2026-02-20"
151 }
152 }
153}
154```
155
156#### Markdown Summary (also generated)
157```markdown
158# Content Inventory: example.com
159**Crawled:** 2026-02-25 | **Total pages:** 347
160
161## Content by Type
162| Type | Count | % |
163|------|-------|---|
164| Blog Posts | 89 | 25.6% |
165| Landing Pages | 23 | 6.6% |
166| ...
167
168## Content by Topic Cluster
169| Topic | Posts | Most Recent |
170|-------|-------|-------------|
171| Cloud Cost Optimization | 34 | 2026-02-20 |
172| ...
173
174## Publishing Cadence
175- Average: 4.2 posts/month
176- Trend: Increasing (3.1 → 5.4 over last 6 months)
177- Most recent: 2026-02-20
178
179## Full Catalog
180| # | Date | Type | Topic | Title | URL |
181|---|------|------|-------|-------|-----|
182| 1 | 2026-02-20 | blog-post | Cloud Cost | How to Reduce... | https://... |
183```
184
185## Tips
186
187- **Sitemap.xml is the best source.** Most well-maintained sites have one. If missing, it's itself an SEO signal (negative).
188- **RSS only shows recent content.** If you need the full catalog, sitemap is essential. RSS is supplementary.
189- **Deep analysis is optional but valuable.** Use it when feeding into brand-voice-extractor or when you need funnel stage mapping.
190- **JS-rendered sites** may need the Apify fallback. Signs: sitemap.xml returns HTML, or blog page returns mostly JavaScript.
191- **Combine with seo-domain-analyzer** to overlay traffic data on the content inventory — see which content actually performs.
192
193## Dependencies
194
195- Python 3.8+
196- `requests` library (`pip install requests`)
197- `APIFY_API_TOKEN` env var (only for Apify fallback mode)