Scrapy Distributed Crawler Framework

Orchestrates large-scale web crawling using Scrapy with scrapy-redis for distributed job queuing. Integrates Splash for JavaScript rendering, stores results in MongoDB via scrapy-mongodb pipeline, and respects robots.txt with AutoThrottle.

agentskillexchange Updated 28 repo stars

File contents

Scrapy Distributed Crawler Framework

Orchestrates large-scale web crawling using Scrapy with scrapy-redis for distributed job queuing. Integrates Splash for JavaScript rendering, stores results in MongoDB via scrapy-mongodb pipeline, and respects robots.txt with AutoThrottle.

Installation

Use the upstream install or setup path that matches your environment:

  • pip install scrapy

Requirements and caveats from upstream:

  • :alt: Supported Python Versions
  • It is cross-platform, and requires Python 3.10+. It is maintained by Zyte_

Basic usage or getting-started notes:

Source

agentskillexchange/skills/tree/main/skills/scrapy-distributed-crawler-framework commit e4ec021fbe

Frequently asked questions

npx skillmds@latest add agentskillexchange/scrapy-distributed-crawler-framework