Web Scraping via MCP
Use this skill to extract clean, readable content from any URL. Returns markdown text, links, and metadata. Free alternative to Firecrawl.
Available Tools
| Tool | What it does |
|---|---|
scrape_url |
Extract clean text content from a URL (Readability-powered) |
extract_links |
Get all links with href and anchor text |
extract_metadata |
Get title, description, OG tags, canonical, favicon |
search_page |
Search for a query string within the page content |
scrape_multiple |
Batch scrape multiple URLs, get title + excerpt per URL |
Workflow
scrape_urlfor reading a single page (docs, blog post, article)extract_linksto discover linked resources from a pageextract_metadatafor SEO analysis or link preview datascrape_multipleto survey multiple pages at once
Key Patterns
- Uses Mozilla Readability (Firefox Reader View engine); works best with server-rendered content
- Does NOT handle JavaScript-heavy SPAs (React apps, dashboards); use a browser MCP for those
scrape_multiplereturns title + excerpt per URL, not full content; use for surveyingsearch_pagesearches within the extracted content, not raw HTML
Error Scenarios
- Empty or tiny markdown: page may be SPA-only or behind login; try a browser MCP or a direct API doc export
- HTTP 403/401: site blocks bots; no fix inside this server
- Timeout or DNS errors: retry with a simpler URL or check network
scrape_multiplepartial failures: inspect per-URL errors in the batch result and scrape successes individually
Limitations
- No headless browser — won't execute JavaScript
- Best for: documentation, blogs, articles, news, wikis
- Won't work for: login-gated content, SPAs, dynamically loaded content