Results for “web-scraping”
69 skillsscrapingbee-automation
Automate web scraping tasks using Scrapingbee through Composio's Rube MCP toolkit, with tool discovery and connection management.
66.9k
hasdata
Extract public web data, search engine results, and structured data from platforms like Google, Amazon, and Zillow using HasData APIs.
42.4k · bundle
web-scraping
Activates for web scraping and Actor development, discovering APIs via traffic interception, recommending optimal strategies, and implementing iteratively. For production, it guides TypeScript Actor creation via Apify CLI.
0 · bundle
web-access
Handles all network operations including search, web scraping, login-required access, and social media content retrieval via a real browser CDP proxy.
2 · bundle
firecrawl-download
Saves an entire website as local files by mapping pages and scraping each into organized directories, supporting multiple formats and screenshots.
2
web-scraping
Extrae datos de sitios web de forma ética usando requests, BeautifulSoup, Selenium o Playwright, respetando robots.txt y aplicando rate limiting.
0 · bundle
More results
firecrawl-automation
Automate web crawling and data extraction with Firecrawl: scrape pages, crawl sites, extract structured data, batch scrape URLs, and map website structures.
66.9k
page-prep
Detects and removes disruptive overlays (cookie banners, modals, paywalls, login walls) from webpages before screenshots, scraping, or browser automation.
142 · bundle
zenrows-automation
Automates Zenrows web scraping operations through Composio's Zenrows toolkit via Rube MCP, with tool discovery and connection management.
66.9k
hasdata-cli
Provides command-line access to search, scraping, and structured web data from over 40 APIs including Google, Amazon, Yelp, and Zillow.
42.4k · bundle
web-access
Handles all networked operations through a real browser via CDP, including search, page scraping, login-required actions, and social media content extraction.
0 · bundle
hasdata
Extract public web data via HasData APIs, including search engine results, structured data from ecommerce, travel, jobs, and local business platforms, with support for web scraping, pre-parsed APIs, and async jobs.
3 · bundle
scrapling
Route web-scraping work into the lightest workable Scrapling mode instead of defaulting to a browser. Use when the user needs HTML extraction, JS-rendered page retrieval, protected-target escalation, quick CLI scraping, agent-facing MCP access, or a larger crawl with Scrapling spiders. Triggers on: scrapling, scrape website, crawl site, adaptive scraping, selector drift, stealthy fetch, browser scraping, scrape to markdown, scrapling mcp, scrapling spider, research harvesting, literature scraping, paper metadata.
42 · bundle
browser
Automates web browser interactions via the browse CLI, including navigation, form filling, screenshots, and scraping, with support for local and remote Browserbase sessions.
1 · bundle
firecrawl-search
Searches the web and optionally extracts full page content, returning results as JSON files.
2
scrape-webpage
Extract content, metadata, and images from a webpage for import or migration to AEM Edge Delivery Services.
142 · bundle
extract
Crawl a live website, extract its design system, brand surface, and page inventory, and save the snapshot under stardust/current/ for downstream redesign or migration.
142 · bundle
defuddle
Extract clean markdown content from web pages using Defuddle CLI, removing clutter and navigation to save tokens.
42.4k
web-search-scraper-api-skill
Extracts clean Markdown content from any website URL using the BrowserAct Web Search Scraper API, with automatic retry and error handling.
3.7k · bundle
firecrawl-crawl
Bulk extract content from an entire website or site section by crawling pages that follow links, with configurable depth, path filters, and concurrency.
2
firecrawl
Scrape and crawl websites for AI with Firecrawl — scrape single URLs to clean Markdown/HTML, crawl entire sites with depth/path filters, extract structured data with LLM schema, use map to discover all URLs, and batch scrape multiple pages in parallel.
2
identify-page-structure
Analyze scraped webpage content to identify section boundaries and content sequences for AEM Edge Delivery Services import.
142 · bundle
webcrawler-deep-crawl
Deep-crawl any website from start URLs, returning per-page LLM-ready text, markdown, or HTML with metadata and in-scope outbound links.
3.7k · bundle
python-executor
Execute Python code in a safe sandboxed environment with 100+ pre-installed libraries for data processing, web scraping, image manipulation, video creation, 3D model processing, PDF generation, API calls, and automation.
584
crawl4ai
Crawl and extract web content for AI with Crawl4AI — async browser-based crawling with clean Markdown output, CSS/XPath/LLM extraction strategies, chunking, screenshot capture, session reuse for SPAs, and Docker deployment.
2
browser-automation
Automate browser tasks, scrape websites, fill forms, capture screenshots, and extract structured data from web pages using Playwright.
20.4k · bundle
firecrawl-map
Discovers and lists all URLs on a website, with optional search filtering to find specific pages within large sites.
2
crawl4ai-mcp-server
Self-hosted web crawling and content extraction exposed as MCP tools, with depth control and clean markdown output.
28
page-collect
Extract structured resources (icons, metadata, text, forms, videos, social links) from any webpage using playwright-cli.
142 · bundle
defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing clutter and navigation to save tokens. Prefer over WebFetch for reading or analyzing standard web pages.
2
firecrawl
Search the web, scrape pages, crawl sites, and interact with dynamic content via the Firecrawl CLI, returning clean markdown for LLM contexts.
2 · bundle
opencli
Turns websites into CLI commands, reusing an existing Chrome login session for browser-backed data extraction and structured output.
61
webapp-testing
Test local web applications by writing native Python Playwright scripts, with helpers for server lifecycle management and a reconnaissance-then-action pattern for dynamic UIs.
158k · bundle
fetch
Retrieve HTML or JSON from static pages, inspect status codes and headers, follow redirects, or get page source for simple scraping without a full browser session. Supports proxies and redirect control.
3.6k · bundle
agent-browser
Automates Chrome/Chromium via CDP with accessibility-tree snapshots and element refs, covering web pages, Electron apps, Slack, and cloud browsers.
2
stardust
Guided multi-page redesign of an existing website through a four-phase pipeline — extract, direct, prototype, and migrate. Tracks progress incrementally per page so redesigns are resumable.
142 · bundle