Plugins
12 pluginscurated
Azure Data Analytics
For data engineers to query and manage big data on Azure with Kusto and Data Lake.
4 skills · plugin
@om-scogo
Data
Data from om-scogo/skillsh-scraper.
100 skills · plugin
@mukul975-2
Privacy Data Protection Skills
Privacy Data Protection Skills from mukul975/Privacy-Data-Protection-Skills.
100 skills · plugin
@nivkazdan
Data Analysis
Data Analysis from nivkazdan/skills-agents-catalog.
6 skills · plugin
curated
Data & ML
SQL, analytics, datasets, models and machine-learning workflows.
29 skills · plugin
@phuryn
Data Analytics
Data analytics skills for PMs: SQL query generation and cohort analysis. Analyze user data, generate queries, and identify retention patterns.
3 skills · plugin
curated
Python Data Visualization
For data scientists to create static and interactive plots using Python libraries.
12 skills · plugin
curated
Deploy Azure Infrastructure
Creates databases, caches, and configures authentication, monitoring, and backup.
3 skills · plugin
curated
Social Media Scraping
Extract structured data from social media platforms via browser automation.
12 skills · plugin
@atc-net
Azure
Azure services skills covering 200+ cloud services, IoT, AI, data, networking, and more
78 skills · plugin
curated
Build GraphQL API
Design a GraphQL schema, implement resolvers with DataLoader, and integrate with Apollo.
4 skills · plugin
@redpanda-data
Redpanda Data Skills
Agent Skills for Redpanda's five products — Streaming (Kafka-compatible engine), SQL (Oxla), Connect (incl. CDC connectors), Cloud (Serverless, BYOC, Dedicated), and the Agentic Data Plane — plus the rpk CLI. Grounded in Redpanda source, docs, and APIs.
32 skills · plugin
Results for “data”
3,187 skillsnuscenes-a-multimodal-dataset-for-autonomous-driving-arxiv-1
nuScenes: A Multimodal Dataset for Autonomous Driving
6
deduplicating-training-data-makes-language-models-better-arx
Deduplicating Training Data Makes Language Models Better
6
imagenet-a-large-scale-hierarchical-image-database-crossref-
ImageNet: A Large-Scale Hierarchical Image Database
6
data-analysis
Analyze campaign performance data — KPI dashboards, weekly/monthly reports, traffic and lead analysis for any active brand
0 · bundle
clickhouse-io
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
0
data-validation-pipelines
A validation *boundary* is any point where data crosses from a system you do not
2
clickhouse-io
ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
1
pdf-data-extractor
Extracts structured data from PDF documents including tables, forms, and scanned images using OCR
6 · bundle
personal-data-test
Classifies personal vs non-personal data per GDPR Art. 4(1) definition test with decision tree for borderline cases. References Breyer v Germany CJEU C-582/14 dynamic IP ruling and WP29 Opinion 4/2007. Keywords: personal data, GDPR Art 4, data classification, Breyer ruling, identifiability test, PII.
228 · bundle
dataverse-python-advanced-patterns
Generate production-ready Python code for Dataverse SDK with advanced patterns including error handling, batch operations, OData optimization, and Pandas integration.
36.2k
googlebigquery-automation
Run SQL queries, explore datasets and metadata, and execute MBQL queries on Google BigQuery through a Metabase integration using Rube MCP (Composio).
66.9k
polars
Process data with high-performance DataFrames using Polars' expression-based API, lazy evaluation, and parallel execution for ETL, analytics, and pandas migration.
30.2k · bundle
hypogenic
Automates hypothesis generation and testing on tabular datasets using LLMs, combining data-driven discovery with literature integration for scientific research.
30.2k · bundle
exploratory-data-analysis
Automatically detect and analyze scientific data files across 200+ formats, generating detailed markdown reports with quality metrics and analysis recommendations.
30.2k · bundle
dask
Scales pandas and NumPy workflows to datasets larger than memory using parallel and distributed computing, with support for dataframes, arrays, bags, and custom task graphs.
253 · bundle
datadog-logs
Query and filter Datadog logs from the shell using the Composio CLI. Run scoped log searches, pivot across services/environments, and export structured JSON for downstream agents instead of click-driving the Datadog UI.
3
data-labeling
Set up and manage data labeling workflows using manual annotation tools, semi-automated pipelines, active learning, and programmatic weak supervision. Use when the user requests data labeling or provides relevant inputs for this workflow.
159
shuffle-json-data
Shuffle repetitive JSON objects safely by validating schema consistency before randomising entries.
36.2k
agent-data-quality
Data Quality Specialist IA — Expert en qualité des données (profiling, cleaning, déduplication, validation de schéma, Great Expectations)
6
dbt
Transforms data in warehouses using dbt with SQL models, tests, and documentation. Use for analytics engineering and data transformation.
2 · bundle
weaviate
Deploys Weaviate vector database with hybrid search, modules, and GraphQL API.
2 · bundle
data-manager-api-audience-ingestion
Uploads audience members to Google products like Customer Match and mobile device ID audiences using the Data Manager API.
14.4k
amc-run-sample-calibration
Run end-to-end calibration on the bundled sample dataset against a running AMC microservice to verify the stack works before processing real data.
2.2k · bundle
detecting-s3-data-exfiltration-attempts
Analyze CloudTrail, GuardDuty, Macie, and VPC Flow Logs to detect unauthorized bulk downloads and cross-account data transfers from AWS S3.
24.6k · bundle
database-connections
Connect to PostgreSQL, MySQL, SQLite, and SQL Server databases using SQLAlchemy, Pandas, and DuckDB. Read and write tables, manage sessions, and handle connection strings securely.
0
polars
Fast in-memory DataFrame library for datasets that fit in RAM. Use when pandas is too slow but data still fits in memory. Lazy evaluation, parallel execution, Apache Arrow backend. Best for 1-100GB da
6
azure-cosmos-db
Expert knowledge for Azure Cosmos DB development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations & coding patterns, and deployment. Use when using Cosmos DB NoSQL/Mongo/Cassandra/PostgreSQL APIs, change feed, vector search, multi-region, or AI/RAG workloads, and other Azure Cosmos DB related development tasks. Not for Azure Table Storage (use azure-table-storage), Azure SQL Database (use azure-sql-database), Azure Database for MySQL (use azure-database-mysql), Azure Database for PostgreSQL (use azure-database-postgresql).
3 · bundle
detecting-data-and-model-poisoning
Detect poisoned training data and backdoored models across the ML pipeline using statistical analysis, activation clustering, and spectral signatures.
24.6k · bundle
studio-export
Exports a platform feature as a product constitution and data reference, including data model, API contracts, and acceptance criteria, for use in AI prototyping tools.
1
managing-neon
Manages Neon serverless Postgres databases via the neonctl CLI and Neon API, covering projects, branches, databases, roles, endpoints, and compute scaling with a discovery-first workflow.
7
aya-eval
Evaluates open-ended generation quality of multilingual LLMs across brainstorming, planning, and long-form tasks, using AYA and DOLLY datasets with qualitative fluency and quality scoring.
3
vcdpa-compliance
Virginia Consumer Data Protection Act (VCDPA) compliance implementation. Covers 5 consumer rights, controller obligations, processor requirements, opt-in for sensitive data, data protection impact assessments, AG enforcement, and cure period provisions. Effective January 1, 2023.
228 · bundle
onchain-data-analytics
Use this skill for on-chain data, explorers, Dune-style queries, wallets, transfers, contract events. Trigger when the task involves crypto work related to Onchain Data Analytics, production implementation, audits, debugging, strategy, or validation.
1 · bundle
prediction-market-data
Access prediction market data from Polymarket and Kalshi, including markets, prices, trades, orderbooks, positions, and cross-platform market matching. Use when you need current odds, historical market data, or wallet-level prediction market analysis.
1 · bundle
whoop
Fetches latest WHOOP recovery, sleep, and strain data and generates daily suggestions.
10 · bundle
docvqa-a-dataset-for-vqa-on-document-images-arxiv-2007-00398
DocVQA: A Dataset for VQA on Document Images
6