Results for “web-scraping”
157 skillsScrape Webpage
Extract content, metadata, and images from a webpage for import or migration to AEM Edge Delivery Services.
142 · bundle
Defuddle
Extract clean markdown content from web pages using Defuddle CLI, removing clutter and navigation to save tokens.
42.4k
Synapse Web Scrape
Fetches and extracts readable text from a URL, then summarizes or presents key information with source citation.
14 · bundle
Zhihu Fetcher
抓取知乎收藏夹与文章正文为 Markdown,支持图片本地化、断点续传及写入 Obsidian 知识库。
28 · bundle
Tiktok Hashtag Videos
Extract paginated video lists for a TikTok hashtag, including author profiles, engagement stats, music, and video metadata.
3.7k · bundle
Python Executor
Execute Python code in a safe sandboxed environment with 100+ pre-installed libraries for data processing, web scraping, image manipulation, video creation, 3D model processing, PDF generation, API calls, and automation.
584
Web Search Scraper API Skill
Extracts clean Markdown content from any website URL using the BrowserAct Web Search Scraper API, with automatic retry and error handling.
3.7k · bundle
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing navigation and clutter to reduce token usage.
20
Browser Use
Direct browser control via CDP for web interaction: automation, scraping, testing, screenshots, and site/app work.
0
Page Collect
Extract structured resources (icons, metadata, text, forms, videos, social links) from any webpage using playwright-cli.
142 · bundle
Google News API Skill
Extracts structured news data from Google News via the BrowserAct API, including headlines, sources, publication times, and article links.
3.7k · bundle
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing clutter and navigation to save tokens. Prefer over WebFetch for reading or analyzing standard web pages.
2
Browser Automation
Automate browser tasks, scrape websites, fill forms, capture screenshots, and extract structured data from web pages using Playwright.
20.4k · bundle
Firecrawl
Search the web, scrape pages, crawl sites, and interact with dynamic content via the Firecrawl CLI, returning clean markdown for LLM contexts.
2 · bundle
Tiktok Search Videos
Extract paginated TikTok video search results for a given keyword, including author details, engagement stats, music info, and video metadata.
3.7k
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing navigation and clutter to reduce token usage. Prefer over WebFetch for reading or analyzing standard web pages.
5
Crawl4ai MCP Server
Self-hosted web crawling and content extraction exposed as MCP tools, with depth control and clean markdown output.
28
Opencli
Turns websites into CLI commands, reusing an existing Chrome login session for browser-backed data extraction and structured output.
61
Firecrawl Map
Discovers and lists all URLs on a website, with optional search filtering to find specific pages within large sites.
2
Facebook Page Profile Posts
Extracts posts from public Facebook pages or profiles, returning structured data with text, engagement metrics, reaction breakdowns, media thumbnails, and pagination support.
3.7k · bundle
Firecrawl Scrape
Extracts clean, LLM-optimized markdown from any URL, including JavaScript-rendered SPAs, with support for concurrent scraping of multiple URLs and options like main-content-only extraction and custom output formats.
2
Webapp Testing
Test local web applications by writing native Python Playwright scripts, with helpers for server lifecycle management and a reconnaissance-then-action pattern for dynamic UIs.
158k · bundle
Defuddle
Extracts clean markdown from web pages via the Defuddle CLI, removing navigation and clutter to reduce token usage for reading or analyzing URLs.
3
Identify Page Structure
Analyze scraped webpage content to identify section boundaries and content sequences for AEM Edge Delivery Services import.
142 · bundle
Producthunt Launches
Extract structured product launch data from Product Hunt leaderboards, enriched with maker profiles and website contact information.
3.7k · bundle
Firecrawl Crawl
Bulk extract content from an entire website or site section by crawling pages that follow links, with configurable depth, path filters, and concurrency.
2
Fetch
Retrieve HTML or JSON from static pages, inspect status codes and headers, follow redirects, or get page source for simple scraping without a full browser session. Supports proxies and redirect control.
3.6k · bundle
Firecrawl Agent
Extracts structured JSON data from complex multi-page websites using an AI agent that navigates pages and returns results matching a schema.
2
Agent Browser
Automates Chrome/Chromium via CDP with accessibility-tree snapshots and element refs, covering web pages, Electron apps, Slack, and cloud browsers.
2
Competitor Profiling
Research and profile competitors from their URLs into actionable intelligence.
36.3k · bundle
Stardust
Guided multi-page redesign of an existing website through a four-phase pipeline — extract, direct, prototype, and migrate. Tracks progress incrementally per page so redesigns are resumable.
142 · bundle
Webcrawler Deep Crawl
Deep-crawl any website from start URLs, returning per-page LLM-ready text, markdown, or HTML with metadata and in-scope outbound links.
3.7k · bundle
Deepapi
Scrapes public web data from LinkedIn, Twitter, GitHub, YouTube, and other sources, sends and reads email, performs deep research, generates images, and searches the web through the DeepAPI service.
42.4k
Performing Paste Site Monitoring For Credentials
Monitor paste sites like Pastebin and GitHub Gists for leaked credentials, API keys, and sensitive data using automated scraping and keyword matching to detect breaches early.
24.6k · bundle
Page Reduce
Reduces a webpage to a structural skeleton by tokenizing content in the browser and applying LLM reasoning to collapse repeated patterns.
142 · bundle
Tavily Crawl
Crawls websites and saves content from multiple pages as local markdown files using the Tavily CLI, with options for depth, breadth, path filtering, and semantic focus.
2