# Crawler

> Web crawling and scraping reference — robots.txt protocol, Scrapy framework, anti-bot detection, headless browsers, and legal considerations

- Skill: `bytesagain/crawler` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add bytesagain/crawler`
- Raw SKILL.md: https://api.skillmd.com/api/skills/bytesagain/crawler/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: bytesagain (https://skillmd.com/u/bytesagain)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/bytesagain/crawler

---


# Crawler

Web crawling and scraping reference — robots.txt protocol, Scrapy framework, anti-bot detection, headless browsers, and legal considerations. No API keys or credentials required — outputs reference documentation only.

## Commands

| Command | Description |
|---------|-------------|
| `intro` | Crawling vs scraping, robots.txt, sitemap |
| `standards` | HTTP caching, structured data, meta tags |
| `troubleshooting` | Anti-bot detection, JS rendering, encoding |
| `performance` | Concurrency, dedup, incremental, distributed |
| `security` | Legal landscape, ethical guidelines, proxies |
| `migration` | BeautifulSoup to Scrapy, requests to Playwright |
| `cheatsheet` | Scrapy commands, CSS/XPath, curl, user-agents |
| `faq` | Legality, JS pages, blocking, storage |

## Output Format

All commands output plain-text reference documentation via heredoc. No external API calls, no credentials needed, no network access.

---

*Powered by BytesAgain | bytesagain.com | hello@bytesagain.com*

