Web Scraping

Web scraping for RAG ingestion. Covers Firecrawl API, Crawl4AI (open source), trafilatura (main-content extraction), BeautifulSoup + readability-lxml, Scrapy for large crawls, Playwright for JS-rendered pages, sitemap.xml discovery, robots.txt, polite crawling, anti-bot, and deduplication. USE WHEN: user mentions "web scraping", "crawl website", "Firecrawl", "Crawl4AI", "trafilatura", "readability", "BeautifulSoup", "Scrapy", "Playwright scrape", "sitemap", "robots.txt", "crawl for RAG", "extract article" DO NOT USE FOR: parsing already-downloaded HTML files for layout - use `unstructured-io`; PDF downloads from the web - use `pdf-extraction`; office docs behind login - use `office-docs`; structured markdown vaults - use `markdown-structured`

claude-dev-suite Updated 28 repo stars

File contents

claude-dev-suite/claude-dev-suite/tree/main/skills/document-processing/web-scraping commit 8b28fd668d

Frequently asked questions

npx skillmds@latest add claude-dev-suite/web-scraping