Content Inventory Tool
Transforms Screaming Frog CSV exports into client-ready audit spreadsheets with two pipelines: pages and files.
Setup
The tool is bundled with this skill at ${CLAUDE_SKILL_DIR}/scripts/. Before first run, install dependencies:
pip install pandas openpyxl
# Only needed if using --follow-redirects:
pip install requests
Quick Start
python ${CLAUDE_SKILL_DIR}/scripts/run_inventory.py \
--pages <raw-pages.csv> \
--orphans <orphan-pages.csv> \
--files <raw-files.csv> \
--inlinks <inlinks.csv> \
--domain example.gov \
--prefix CLIENT \
--output-dir output \
--format xlsx
This produces:
output/CLIENT-audit-all-pages.xlsx— pages inventoryoutput/CLIENT-audit-all-files.xlsx— files inventory
For CSV output, omit --format or use --format csv.
Input Files
Four Screaming Frog exports are required. See screaming-frog-exports.md for detailed SF configuration instructions.
| Flag | SF Export | Columns the tool uses |
|---|---|---|
--pages |
Internal > HTML (all pages crawl export) | Address, Title 1, Status Code, Flesch Reading Ease Score, GA4 Views, Redirect URL |
--orphans |
Crawl Analysis > Orphan Pages | URL (lacks https:// prefix — this is expected) |
--files |
Internal > All (non-HTML resources) | Address, Title 1, Status Code, Size (bytes) |
--inlinks |
Bulk Export > All Inlinks | From, To, Anchor Text, Alt Text, Size |
The raw pages CSV contains ~79 columns; the tool extracts only 6. Extra columns are ignored.
Options
| Flag | Default | Description |
|---|---|---|
--domain |
(none) | Expected domain (e.g., energy.maryland.gov). Flags rogue off-domain URLs in output. |
--follow-redirects |
off | Follow HTTP redirect chains against the live site. Adds ~0.2s per URL. Requires requests. |
--format |
csv |
Output format: csv or xlsx. |
--output-dir |
. |
Directory for output files. Created if it doesn't exist. |
--prefix |
output |
Filename prefix for output files. |
What the Tool Does
Pages pipeline
- Load raw pages, extract 6 columns, normalize URLs
- Remove 404 status pages
- Append orphan pages (separate CSV, URLs already lack
https://) - Deduplicate on normalized URL — first non-empty title/status, MAX for GA views and reading scores
- Filter out non-page URLs (PDFs, docs, images, etc.)
- Resolve redirects from Screaming Frog's Redirect URL column (or via HTTP with
--follow-redirects) - Merge redirect duplicates — multiple source URLs become newline-separated in one cell
- Add redirect target URLs not already in inventory
- Flag rogue URLs outside
--domain - Convert Flesch reading scores to grade levels (blank stays blank, not "5th grade")
Files pipeline
- Load raw files + inlinks, normalize URLs
- Filter out page URLs (.aspx, .html, .htm, .php, .jsp)
- Collect unique file URLs from both sources
- Enrich each file with: title, type (classified by extension), anchor text, source page URLs (newline-separated), alt text, size
Understanding the Output
Both output files use a two-header-row format:
- Row 1: Category headers (e.g., "Page information", "Review criteria", "Decisions & comments")
- Row 2: Column descriptions with embedded newlines explaining each field
Pre-filled data columns are followed by empty review columns for auditors. The "Required by law" column defaults to FALSE.
For complete column specifications, see output-columns.md.
Pre-run Validation
Before running the full tool, validate that input CSVs have the expected columns:
python ${CLAUDE_SKILL_DIR}/scripts/check_inputs.py \
--pages <raw-pages.csv> \
--orphans <orphan-pages.csv> \
--files <raw-files.csv> \
--inlinks <inlinks.csv>
This catches the most common error: passing the wrong Screaming Frog export to the wrong flag.
Troubleshooting
Common issues:
- KeyError on a column name — Wrong SF export passed to wrong flag. Run
check_inputs.pyto diagnose. - ModuleNotFoundError: pandas/openpyxl — Run
pip install pandas openpyxl. - Output has 0 rows — All pages were 404, or the wrong file was passed to
--pages.
For more, see troubleshooting.md.
Development
# Run all tests
pytest
# Run a single test
pytest tests/test_normalize.py::test_strips_https
Source modules in content_inventory/: cli.py (arg parsing), pages.py (pages pipeline), files.py (files pipeline), output.py (formatting), normalize.py (URL normalization), redirects.py (redirect resolution), filetypes.py (file classification), readability.py (Flesch to grade).