Results for “extraction”
70 skillsDocument AI
Comprehensive patterns for AI-powered document understanding including PDF parsing, OCR, invoice/receipt extraction, table extraction, multimodal RAG with vision models, and structured data output. Use when "document parsing, PDF extraction, OCR, invoice processing, receipt extraction, document understanding, LlamaParse, Unstructured, vision document, table extraction, structured output from PDF, " mentioned.
128 · bundle
Agent Web Scraper
Web Scraper IA — Expert en extraction web (Scrapy, BeautifulSoup, Playwright, anti-bot, proxy rotation, data extraction)
6
Web Scraper
Web scraping and content comprehension agent — multi-strategy extraction with cascade fallback, news detection, boilerplate removal, structured metadata, and LLM entity extraction
228 · bundle
Taggun Automation
Automate Taggun document data extraction operations through Composio's Taggun toolkit via Rube MCP.
66.9k
Diffbot Automation
Automate Diffbot data extraction and analysis through Composio's Diffbot toolkit via Rube MCP.
66.9k
Dynamic Workflow Mode
Design task-local harnesses, eval gates, and reusable skill extraction for adaptive agent workflows.
226k
More results
Crawl4ai MCP Server
Self-hosted web crawling and content extraction exposed as MCP tools, with depth control and clean markdown output.
28
Tavily Best Practices
Reference for building Tavily-powered search, extraction, crawling, and research into agentic workflows and RAG systems.
2 · bundle
Scrape Do Automation
Automate web scraping and data extraction tasks using the Scrape Do toolkit via Rube MCP and Composio.
66.9k
Hwp
Use kordoc for agent-native HWP/HWPX document parsing, JSON extraction, diffing, form-field extraction, and Markdown→HWPX reverse conversion (read/convert only — for binary editing use rhwp-edit).
3 · bundle
Parsehub Automation
Automates Parsehub data extraction tasks through Composio's Parsehub toolkit via Rube MCP, with tool discovery and connection management.
66.9k
Hwp
Use kordoc for agent-native HWP/HWPX document parsing, JSON extraction, diffing, form-field extraction, and Markdown→HWPX reverse conversion (read/convert only — for binary editing use rhwp-edit).
0 · bundle
200 Aeon E7807df1
Guides feature extraction and preprocessing for time series data using aeon transformers, covering collection and series transformers with code examples.
7 · bundle
Mini Context Graph
Build a persistent, compounding knowledge base that combines a wiki, knowledge graph, and raw source storage for structured retrieval with provenance.
36.2k · bundle
Malware Analysis
Analyze suspected malware through static, dynamic, and behavioral techniques, including IOC extraction, YARA or Sigma rules, sandboxing, and anti-analysis behavior detection.
12.8k · bundle
Tavily Web
Web search, content extraction, crawling, and research capabilities using Tavily API
505
Ocr
从OCR识别后的医疗票据文本中提取日期、医生姓名、病人姓名、诊断和总消费,并进行文本矫正,输出JSON格式。
559
Agent Songsee V2
Expert en analyse audio avancé (spectrograms, mel, chroma, MFCC, feature extraction, CLI)
6
Board Meeting
Runs a structured 6-phase multi-agent board meeting protocol for strategic decisions, with isolated C-suite contributions, critic analysis, synthesis, founder review, and decision extraction.
20.4k · bundle
Histolab
Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.
0 · bundle
Detecting Model Extraction Attacks
Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability.
24.6k · bundle
Data Scraping
Builds a configurable scraping agent that collects data from APIs, HTML, or RSS, enriches it with Gemini AI scoring, and stores results in Notion, Google Sheets, Supabase, or local files.
1 · bundle
Browser Use
Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, or extract information from web pages.
1
Graph RAG
Knowledge-graph-augmented retrieval. Entity and triple extraction, graph construction (Neo4j, LlamaIndex PropertyGraphIndex), hierarchical community summarization (Microsoft GraphRAG), personalized PageRank (HippoRAG), multi-hop traversal retrieval, and hybrid graph + vector pipelines. USE WHEN: user mentions "GraphRAG", "HippoRAG", "knowledge graph RAG", "entity extraction", "multi-hop reasoning", "Neo4j RAG", "LlamaIndex property graph", "LangChain graph retriever", "triple extraction", "community summarization" DO NOT USE FOR: vanilla vector RAG - use `rag-patterns`; multimodal inputs - use `multimodal-rag`; production indexing ops - use `rag-production`; hallucination checks - use `rag-guardrails`
28
Instructor
Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library
0 · bundle
Instructor
Extract structured data from LLM responses with Pydantic validation, retry failed extractions automatically, parse complex JSON with type safety, and stream partial results with Instructor - battle-tested structured output library
1 · bundle
Playwright
Use when the task requires automating a real browser from the terminal (navigation, form filling, snapshots, screenshots, data extraction, UI-flow debugging) via `playwright-cli` or the bundled wrapper script.
0 · bundle
Web Search
Search the web and extract content from URLs using Tavily and Exa APIs via the inference.sh CLI.
584
Pharo Refactor
Refactor messy Pharo code safely through the genie MCP tools, working from the live image (no files). Use for renames, method extraction, moving behavior, or general cleanup of existing code.
15
Youtube Search API Skill
Extracts structured data from YouTube search results, including videos, shorts, channels, and playlists, using the BrowserAct API.
3.7k · bundle
Wechat Article Search API Skill
Extract full article contents from WeChat using the BrowserAct API, with keyword search and date filtering.
3.7k · bundle
Histolab
Process whole slide images for digital pathology: detect tissue, extract tiles, and prepare datasets for deep learning pipelines.
30.2k · bundle
Book Sft Pipeline
Convert books into supervised fine-tuning datasets and train style-transfer models that replicate an author's voice.
16.9k · bundle
Tavily Extract
Extracts clean markdown or text from one or more URLs using the Tavily CLI, with support for JavaScript-rendered pages and query-focused chunking.
2
Firecrawl Agent
Extracts structured JSON data from complex multi-page websites using an AI agent that navigates pages and returns results matching a schema.
2
Youtube Batch Transcript Extractor API Skill
Extracts YouTube video transcripts and metadata in batch via the BrowserAct API, using search keywords and date filters.
3.7k · bundle