Results for “extraction”

70 skills
More results
agentskillexchange
Crawl4ai MCP Server
Self-hosted web crawling and content extraction exposed as MCP tools, with depth control and clean markdown output.
28
scoheart
Tavily Best Practices
Reference for building Tavily-powered search, extraction, crawling, and research into agentic workflows and RAG systems.
2 · bundle
composiohq
Scrape Do Automation
Automate web scraping and data extraction tasks using the Scrape Do toolkit via Rube MCP and Composio.
66.9k
aibot88
Hwp
Use kordoc for agent-native HWP/HWPX document parsing, JSON extraction, diffing, form-field extraction, and Markdown→HWPX reverse conversion (read/convert only — for binary editing use rhwp-edit).
3 · bundle
composiohq
Parsehub Automation
Automates Parsehub data extraction tasks through Composio's Parsehub toolkit via Rube MCP, with tool discovery and connection management.
66.9k
om-scogo
Hwp
Use kordoc for agent-native HWP/HWPX document parsing, JSON extraction, diffing, form-field extraction, and Markdown→HWPX reverse conversion (read/convert only — for binary editing use rhwp-edit).
0 · bundle
tools-only
200 Aeon E7807df1
Guides feature extraction and preprocessing for time series data using aeon transformers, covering collection and series transformers with code examples.
7 · bundle
github
Mini Context Graph
Build a persistent, compounding knowledge base that combines a wiki, knowledge graph, and raw source storage for structured retrieval with provenance.
36.2k · bundle
zhaoxuya520
Malware Analysis
Analyze suspected malware through static, dynamic, and behavioral techniques, including IOC extraction, YARA or Sigma rules, sandboxing, and anti-analysis behavior detection.
12.8k · bundle
dokhacgiakhoa
Tavily Web
Web search, content extraction, crawling, and research capabilities using Tavily API
505
ecnu-icalk
Ocr
从OCR识别后的医疗票据文本中提取日期、医生姓名、病人姓名、诊断和总消费,并进行文本矫正,输出JSON格式。
559
ziri22
Agent Songsee V2
Expert en analyse audio avancé (spectrograms, mel, chroma, MFCC, feature extraction, CLI)
6
alirezarezvani
Board Meeting
Runs a structured 6-phase multi-agent board meeting protocol for strategic decisions, with isolated C-suite contributions, critic analysis, synthesis, founder review, and decision extraction.
20.4k · bundle
michaelschecht
Histolab
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
0 · bundle
mukul975
Detecting Model Extraction Attacks
Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability.
24.6k · bundle
auto-skiller
Data Scraping
Builds a configurable scraping agent that collects data from APIs, HTML, or RSS, enriches it with Gemini AI scoring, and stores results in Notion, Google Sheets, Supabase, or local files.
1 · bundle
neuralblitz
Browser Use
Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.
1
claude-dev-suite
Graph RAG
Knowledge-graph-augmented retrieval. Entity and triple extraction, graph construction (Neo4j, LlamaIndex PropertyGraphIndex), hierarchical community summarization (Microsoft GraphRAG), personalized PageRank (HippoRAG), multi-hop traversal retrieval, and hybrid graph + vector pipelines. USE WHEN: user mentions "GraphRAG", "HippoRAG", "knowledge graph RAG", "entity extraction", "multi-hop reasoning", "Neo4j RAG", "LlamaIndex property graph", "LangChain graph retriever", "triple extraction", "community summarization" DO NOT USE FOR: vanilla vector RAG - use `rag-patterns`; multimodal inputs - use `multimodal-rag`; production indexing ops - use `rag-production`; hallucination checks - use `rag-guardrails`
28
qcmuu
Instructor
Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library
0 · bundle
tianhao909
Instructor
Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library
1 · bundle
michaelschecht
Playwright
Use when the task requires automating a real browser from the terminal (navigation, form filling, snapshots, screenshots, data extraction, UI-flow debugging) via `playwright-cli` or the bundled wrapper script.
0 · bundle
inference-sh
Web Search
Search the web and extract content from URLs using Tavily and Exa APIs via the inference.sh CLI.
584
kentbeck
Pharo Refactor
Refactor messy Pharo code safely through the genie MCP tools, working from the live image (no files). Use for renames, method extraction, moving behavior, or general cleanup of existing code.
15
browser-act
Youtube Search API Skill
Extracts structured data from YouTube search results, including videos, shorts, channels, and playlists, using the BrowserAct API.
3.7k · bundle
browser-act
Wechat Article Search API Skill
Extract full article contents from WeChat using the BrowserAct API, with keyword search and date filtering.
3.7k · bundle
k-dense-ai
Histolab
Process whole slide images for digital pathology: detect tissue, extract tiles, and prepare datasets for deep learning pipelines.
30.2k · bundle
muratcankoylan
Book Sft Pipeline
Convert books into supervised fine-tuning datasets and train style-transfer models that replicate an author's voice.
16.9k · bundle
scoheart
Tavily Extract
Extracts clean markdown or text from one or more URLs using the Tavily CLI, with support for JavaScript-rendered pages and query-focused chunking.
2
scoheart
Firecrawl Agent
Extracts structured JSON data from complex multi-page websites using an AI agent that navigates pages and returns results matching a schema.
2
browser-act
Youtube Batch Transcript Extractor API Skill
Extracts YouTube video transcripts and metadata in batch via the BrowserAct API, using search keywords and date filters.
3.7k · bundle