Results for “web-scraping”

157 skills
adobe
Scrape Webpage
Extract content, metadata, and images from a webpage for import or migration to AEM Edge Delivery Services.
142 · bundle
antigravity
Defuddle
Extract clean markdown content from web pages using Defuddle CLI, removing clutter and navigation to save tokens.
42.4k
upayanghosh
Synapse Web Scrape
Fetches and extracts readable text from a URL, then summarizes or presents key information with source citation.
14 · bundle
handsomestwei
Zhihu Fetcher
抓取知乎收藏夹与文章正文为 Markdown,支持图片本地化、断点续传及写入 Obsidian 知识库。
28 · bundle
browser-act
Tiktok Hashtag Videos
Extract paginated video lists for a TikTok hashtag, including author profiles, engagement stats, music, and video metadata.
3.7k · bundle
inference-sh
Python Executor
Execute Python code in a safe sandboxed environment with 100+ pre-installed libraries for data processing, web scraping, image manipulation, video creation, 3D model processing, PDF generation, API calls, and automation.
584
browser-act
Web Search Scraper API Skill
Extracts clean Markdown content from any website URL using the BrowserAct Web Search Scraper API, with automatic retry and error handling.
3.7k · bundle
galyarderlabs
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing navigation and clutter to reduce token usage.
20
lovits
Browser Use
Direct browser control via CDP for web interaction: automation, scraping, testing, screenshots, and site/app work.
0
adobe
Page Collect
Extract structured resources (icons, metadata, text, forms, videos, social links) from any webpage using playwright-cli.
142 · bundle
browser-act
Google News API Skill
Extracts structured news data from Google News via the BrowserAct API, including headlines, sources, publication times, and article links.
3.7k · bundle
nimoqup046-collab
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing clutter and navigation to save tokens. Prefer over WebFetch for reading or analyzing standard web pages.
2
alirezarezvani
Browser Automation
Automate browser tasks, scrape websites, fill forms, capture screenshots, and extract structured data from web pages using Playwright.
20.4k · bundle
scoheart
Firecrawl
Search the web, scrape pages, crawl sites, and interact with dynamic content via the Firecrawl CLI, returning clean markdown for LLM contexts.
2 · bundle
browser-act
Tiktok Search Videos
Extract paginated TikTok video search results for a given keyword, including author details, engagement stats, music info, and video metadata.
3.7k
lucaspmarie-a11y
Defuddle
Extracts clean markdown content from web pages using the Defuddle CLI, removing navigation and clutter to reduce token usage. Prefer over WebFetch for reading or analyzing standard web pages.
5
agentskillexchange
Crawl4ai MCP Server
Self-hosted web crawling and content extraction exposed as MCP tools, with depth control and clean markdown output.
28
comeonoliver
Opencli
Turns websites into CLI commands, reusing an existing Chrome login session for browser-backed data extraction and structured output.
61
scoheart
Firecrawl Map
Discovers and lists all URLs on a website, with optional search filtering to find specific pages within large sites.
2
browser-act
Facebook Page Profile Posts
Extracts posts from public Facebook pages or profiles, returning structured data with text, engagement metrics, reaction breakdowns, media thumbnails, and pagination support.
3.7k · bundle
scoheart
Firecrawl Scrape
Extracts clean, LLM-optimized markdown from any URL, including JavaScript-rendered SPAs, with support for concurrent scraping of multiple URLs and options like main-content-only extraction and custom output formats.
2
anthropic
Webapp Testing
Test local web applications by writing native Python Playwright scripts, with helpers for server lifecycle management and a reconnaissance-then-action pattern for dynamic UIs.
158k · bundle
phoroth
Defuddle
Extracts clean markdown from web pages via the Defuddle CLI, removing navigation and clutter to reduce token usage for reading or analyzing URLs.
3
adobe
Identify Page Structure
Analyze scraped webpage content to identify section boundaries and content sequences for AEM Edge Delivery Services import.
142 · bundle
browser-act
Producthunt Launches
Extract structured product launch data from Product Hunt leaderboards, enriched with maker profiles and website contact information.
3.7k · bundle
scoheart
Firecrawl Crawl
Bulk extract content from an entire website or site section by crawling pages that follow links, with configurable depth, path filters, and concurrency.
2
browserbase
Fetch
Retrieve HTML or JSON from static pages, inspect status codes and headers, follow redirects, or get page source for simple scraping without a full browser session. Supports proxies and redirect control.
3.6k · bundle
scoheart
Firecrawl Agent
Extracts structured JSON data from complex multi-page websites using an AI agent that navigates pages and returns results matching a schema.
2
samuelpatro
Agent Browser
Automates Chrome/Chromium via CDP with accessibility-tree snapshots and element refs, covering web pages, Electron apps, Slack, and cloud browsers.
2
coreyhaines31
Competitor Profiling
Research and profile competitors from their URLs into actionable intelligence.
36.3k · bundle
adobe
Stardust
Guided multi-page redesign of an existing website through a four-phase pipeline — extract, direct, prototype, and migrate. Tracks progress incrementally per page so redesigns are resumable.
142 · bundle
browser-act
Webcrawler Deep Crawl
Deep-crawl any website from start URLs, returning per-page LLM-ready text, markdown, or HTML with metadata and in-scope outbound links.
3.7k · bundle
antigravity
Deepapi
Scrapes public web data from LinkedIn, Twitter, GitHub, YouTube, and other sources, sends and reads email, performs deep research, generates images, and searches the web through the DeepAPI service.
42.4k
mukul975
Performing Paste Site Monitoring For Credentials
Monitor paste sites like Pastebin and GitHub Gists for leaked credentials, API keys, and sensitive data using automated scraping and keyword matching to detect breaches early.
24.6k · bundle
adobe
Page Reduce
Reduces a webpage to a structural skeleton by tokenizing content in the browser and applying LLM reasoning to collapse repeated patterns.
142 · bundle
scoheart
Tavily Crawl
Crawls websites and saves content from multiple pages as local markdown files using the Tavily CLI, with options for depth, breadth, path filtering, and semantic focus.
2