wigolo crawl
Crawl sites with configurable strategy, depth, and rate limiting. All pages enter the local cache with embeddings.
Quick Reference
// Crawl docs via sitemap (fastest, recommended for doc sites)
{ "url": "https://docs.example.com", "strategy": "sitemap", "max_pages": 30 }
// BFS crawl with scope filter
{ "url": "https://example.com", "strategy": "bfs", "max_depth": 3, "max_pages": 50, "include_patterns": ["^https://example\\.com/docs"] }
// URL discovery only (no content fetched — fastest for scoping)
{ "url": "https://example.com", "strategy": "map" }
// Authenticated crawl
{ "url": "https://app.example.com/docs", "strategy": "bfs", "use_auth": true, "max_pages": 20 }
Parameters
| Parameter | Type | Default | When to use |
|---|---|---|---|
url |
string | required | Seed URL |
strategy |
string | "bfs" | "sitemap" for doc sites, "map" for URL discovery only |
max_depth |
number | 2 | How many link levels to follow |
max_pages |
number | 20 | Hard cap on pages fetched |
include_patterns |
string[] | none | Regex whitelist — ALWAYS add to stay in scope |
exclude_patterns |
string[] | none | Regex blacklist |
use_auth |
boolean | false | For authenticated sites |
extract_links |
boolean | false | Return inter-page link graph |
max_total_chars |
number | 100000 | Total char budget |
max_tokens_out |
number | none | Token-budget cap (cl100k-base) |
include_full_markdown |
boolean | false | Pages return evidence-only by default |
citation_format |
string | "numbered" | "numbered" / "json" / "anthropic_tags" |
Pages are dedup'd by canonical URL (anchor-fragment aware: /intro and /intro#install collapse to one entry).
After Crawling
All crawled pages enter the local cache with embeddings. This means:
cache({ query: "..." })finds content instantly (no network).find_similar({ url: "..." })discovers related pages from cached content.- Future searches that hit cached URLs return instantly.
Crawl first, then use cache and find_similar for subsequent lookups.
Anti-Patterns
- DON'T crawl
max_pages: 100withoutinclude_patterns— fetches nav/footer/sitemap noise. - DON'T use BFS on large doc sites —
strategy: "sitemap"is faster and more complete. - DON'T crawl when you need one page — use
fetch.
When NOT to use wigolo-crawl
- Login-required crawl beyond what
use_authcovers — handle authentication externally first, thencrawl.
See Also
- wigolo-fetch — for single pages
- wigolo-find-similar — discover related content after crawling
- wigolo-cache — query the pages a crawl populated