ETL & Pipelines Agent Skills
ETL & Pipelines
153 skillstoken-scam-analysis
Perform forensic on-chain analysis of EVM tokens to detect scams, rug pulls, and soft rugs by cross-referencing on-chain state against team narratives.
1.2k · bundle
datagma-automation
Automate Datagma operations through Composio's Datagma toolkit via Rube MCP, with tool discovery and connection management.
66.9k
parseur-automation
Automate Parseur document parsing operations through Composio's Parseur toolkit via Rube MCP.
66.9k
listclean-automation
Automates Listclean data-cleaning operations through Composio's toolkit via Rube MCP, with tool discovery and connection management.
66.9k
scrape-do-automation
Automate web scraping and data extraction tasks using the Scrape Do toolkit via Rube MCP and Composio.
66.9k
brightdata-automation
Automate Brightdata web scraping and data collection operations through Composio's Brightdata toolkit via Rube MCP.
66.9k
stormglass-io-automation
Automate Stormglass IO operations through Composio's Stormglass IO toolkit via Rube MCP, including tool discovery, connection management, and execution.
66.9k
browser-act-skill-forge
Turns any website's data extraction or operation needs into reusable Agent-callable Skill packages by exploring API endpoints or DOM methods, then generating SKILL.md and Python scripts.
3.7k · bundle
webcrawler-deep-crawl
Deep-crawl any website from start URLs, returning per-page LLM-ready text, markdown, or HTML with metadata and in-scope outbound links.
3.7k · bundle
amazon-asin-lookup-api-skill
Extract structured product details from Amazon using an ASIN, including title, price, ratings, brand, and description via the BrowserAct API.
3.7k · bundle
dask
Scale pandas and NumPy workflows to larger-than-memory datasets using parallel and distributed computing.
30.2k · bundle
vaex
Process and analyze large tabular datasets (billions of rows) that exceed available RAM using lazy, out-of-core DataFrames with fast aggregations, visualization, and machine learning integration.
30.2k · bundle
polars
Process data with high-performance DataFrames using Polars' expression-based API, lazy evaluation, and parallel execution for ETL, analytics, and pandas migration.
30.2k · bundle
nextflow
Build, run, and debug Nextflow data pipelines and nf-core workflows end to end, covering processes, channels, operators, configuration, testing, and deployment to HPC or cloud.
30.2k · bundle
pacsomatic
Validates inputs, generates samplesheets and launch scripts, and optionally executes nf-core/pacsomatic matched tumor-normal workflows from BAM files, supporting local runs and scheduler submission (LSF/Slurm/PBS/SGE).
30.2k · bundle
pylabrobot
Control liquid handling robots, plate readers, pumps, and other lab equipment through a unified Python interface across platforms.
30.2k · bundle
bioservices
Query 40+ bioinformatics services (UniProt, KEGG, ChEMBL, Reactome) with a unified Python interface for cross-database analysis, identifier mapping, and sequence analysis.
30.2k · bundle
bulk-rnaseq
Orchestrates a complete bulk RNA-seq differential-expression study from raw FASTQ reads through QC, alignment, quantification, differential expression, pathway enrichment, and publication figures.
30.2k · bundle
zarr-python
Store and process large N-dimensional arrays with chunking, compression, and parallel I/O, integrating with NumPy, Dask, and Xarray for cloud-native scientific computing.
30.2k · bundle
dnanexus-integration
Build and deploy apps/applets on the DNAnexus cloud genomics platform, manage data objects, run workflows, and use the dxpy Python SDK for genomics pipeline development and execution.
30.2k · bundle
latchbio-integration
Build and deploy bioinformatics workflows as serverless pipelines on the Latch platform using Python decorators, cloud data management, and GPU support.
30.2k · bundle
benchling-integration
Integrate with Benchling's Python SDK and REST API to manage registry entities, inventory, ELN entries, workflows, and Data Warehouse queries for life sciences R&D automation.
30.2k · bundle
opentrons-integration
Write Opentrons Protocol API v2 protocols for Flex and OT-2 robots to automate liquid handling, control hardware modules, and manage labware configurations.
30.2k · bundle
labarchive-integration
Access and manage LabArchives electronic lab notebooks programmatically via REST API. Create entries, upload attachments, backup notebooks, generate reports, and integrate with Protocols.io, Jupyter, REDCap, and other scientific tools.
30.2k · bundle
processing-stix-taxii-feeds
Processes STIX 2.1 threat intelligence bundles from TAXII 2.1 servers, normalizing objects into platform-native schemas and routing them to consuming systems.
24.6k · bundle
analyzing-threat-intelligence-feeds
Ingests, normalizes, and enriches structured and unstructured threat intelligence feeds into STIX 2.1 format, evaluating feed quality and deduplicating indicators for distribution to SIEM, firewall, and EDR platforms.
24.6k · bundle
performing-ioc-enrichment-automation
Automates multi-source enrichment of IPs, domains, URLs, and file hashes using VirusTotal, AbuseIPDB, Shodan, GreyNoise, URLScan.io, and MISP to provide contextual risk scoring and disposition recommendations for SOC analysts.
24.6k · bundle
implementing-stix-taxii-feed-integration
Consume and produce STIX/TAXII 2.1 cyber threat intelligence feeds using Python, including server discovery, collection polling, object parsing, and SIEM/TIP integration.
24.6k · bundle
implementing-siem-correlation-rules-for-apt
Detect APT lateral movement by chaining Windows authentication events, process execution telemetry, and network connection logs across hosts using Splunk SPL and Sigma rule format.
24.6k · bundle
ml-pipeline
Designs and implements production-grade ML pipeline infrastructure: configures experiment tracking, creates orchestration DAGs, builds feature store schemas, deploys model registries, and automates retraining and validation workflows.
10.4k · bundle
book-sft-pipeline
Convert books into supervised fine-tuning datasets and train style-transfer models that replicate an author's voice.
16.9k · bundle
ray-data
Process large-scale ML datasets with distributed streaming execution across CPU/GPU, supporting Parquet, CSV, JSON, images, and integration with PyTorch, TensorFlow, and Ray Train.
10.4k · bundle
nemo-curator
GPU-accelerated data curation for LLM training, supporting text, image, video, and audio with fuzzy deduplication, quality filtering, semantic deduplication, PII redaction, and NSFW detection.
10.4k · bundle
n8n-binary-and-data
Handle files and binary data in n8n workflows correctly, covering the $binary vs $json split, reading/writing binary, preserving binary across transforms, and the agent-tool binary boundary.
5.7k · bundle
edge-pipeline-orchestrator
Coordinate multi-stage edge research pipelines from candidate detection through strategy design, review, revision, and export.
2.3k · bundle
lets-go-rss
Aggregate RSS feeds from YouTube, Vimeo, Behance, Twitter/X, Bilibili, Weibo, Douyin, Xiaohongshu, and Zhihu with incremental updates, deduplication, and AI classification.
99 · bundle