Packs
12 packscurated
Azure Data Analytics
For data engineers to query and manage big data on Azure with Kusto and Data Lake.
4 skills · pack
@om-scogo
Data
Data from om-scogo/skillsh-scraper.
100 skills · pack
@mukul975-2
Privacy Data Protection Skills
Privacy Data Protection Skills from mukul975/Privacy-Data-Protection-Skills.
100 skills · pack
@nivkazdan
Data Analysis
Data Analysis from nivkazdan/skills-agents-catalog.
6 skills · pack
curated
Data & ML
SQL, analytics, datasets, models and machine-learning workflows.
29 skills · pack
@phuryn
Data Analytics
Data analytics skills for PMs: SQL query generation and cohort analysis. Analyze user data, generate queries, and identify retention patterns.
3 skills · pack
curated
Python Data Visualization
For data scientists to create static and interactive plots using Python libraries.
12 skills · pack
curated
Deploy Azure Infrastructure
Creates databases, caches, and configures authentication, monitoring, and backup.
3 skills · pack
curated
Social Media Scraping
Extract structured data from social media platforms via browser automation.
12 skills · pack
@atc-net
Azure
Azure services skills covering 200+ cloud services, IoT, AI, data, networking, and more
78 skills · pack
curated
Build GraphQL API
Design a GraphQL schema, implement resolvers with DataLoader, and integrate with Apollo.
4 skills · pack
@redpanda-data
Redpanda Data Skills
Agent Skills for Redpanda's five products — Streaming (Kafka-compatible engine), SQL (Oxla), Connect (incl. CDC connectors), Cloud (Serverless, BYOC, Dedicated), and the Agentic Data Plane — plus the rpk CLI. Grounded in Redpanda source, docs, and APIs.
32 skills · pack
Results for “data”
322 skillsclickhouse-io
Provides ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.
226k
data-archive
Documenta, versiona y cierra proyectos de análisis de datos para que queden ordenados y reproducibles en el futuro.
0
tao-validate-dataset-format
Validates NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors using the `tao-daft validate` CLI tool.
2.2k · bundle
data-analysis
Analiza datasets con Pandas y NumPy: explora distribuciones, correlaciones y patrones, y aplica tests de hipótesis para extraer conocimiento no obvio.
0 · bundle
sql-pro
Optimize SQL queries, design database schemas, and tune performance across cloud-native and hybrid OLTP/OLAP environments.
5
pandas-python
Write, review, debug, test, or optimize pandas Series, DataFrame, Index, groupby, merge, reshape, dtype, missing-value, and time-series code.
0 · bundle
google-analytics-data-api-basics
Enables the Google Analytics Data API, authenticates via gcloud, and creates customized reports using the v1beta client library.
14.4k · bundle
cupynumeric-parallel-data-load
Load sharded datasets (npy, Parquet, HDF5, raw binary) into distributed cuPyNumeric arrays using manual partitioning and Legate task launches.
2.2k · bundle
ddia-systems
Design reliable, scalable, and maintainable data systems by applying principles from storage engines, replication, partitioning, transactions, and consistency models.
1.6k · bundle
spark-engineer
Write, optimize, and debug Apache Spark jobs for high-performance distributed data processing, ETL pipelines, and big data workloads.
10.4k · bundle
vaex
Process and analyze tabular datasets larger than RAM using lazy, out-of-core DataFrames, with fast aggregations, visualization, and machine learning integration.
253 · bundle
hasdata
Extract public web data, search engine results, and structured data from platforms like Google, Amazon, and Zillow using HasData APIs.
42.4k · bundle
gov-environment
Fetches real-time EPA air quality data and HUD foreclosure listings through an MCP server, enabling environmental monitoring and housing research.
5
polars
Process in-memory tabular data with a fast, expression-based DataFrame library that supports lazy evaluation, parallel execution, and Apache Arrow semantics.
3
database-lookup
Query documented public database APIs with explicit endpoints, filters, pagination, and provenance for reproducible retrieval of scientific, regulatory, or financial facts.
30.2k · bundle
data-analysis
Guide through a structured data analysis workflow: define the question, validate data quality, select the appropriate analytical method, and produce decision-ready findings with caveats.
42 · bundle
polars
Process data with high-performance DataFrames using Polars' expression-based API, lazy evaluation, and parallel execution for ETL, analytics, and pandas migration.
30.2k · bundle
exploratory-data-analysis
Automatically detect and analyze scientific data files across 200+ formats, generating detailed markdown reports with quality metrics and analysis recommendations.
30.2k · bundle
dask
Scales pandas and NumPy workflows to datasets larger than memory using parallel and distributed computing, with support for dataframes, arrays, bags, and custom task graphs.
253 · bundle
whoop
Fetches latest WHOOP recovery, sleep, and strain data and generates daily suggestions.
10 · bundle
datadog-logs
Query and filter Datadog logs from the shell using the Composio CLI, enabling scoped log searches, pivoting across services and environments, and exporting structured JSON for downstream analysis.
66.9k
jq
Query, filter, transform, and aggregate JSON data using jq in shell pipelines and scripts.
42.4k
power-bi-model-design-review
Evaluates Power BI data model architecture, relationships, storage modes, and performance to identify optimization opportunities and ensure adherence to best practices.
36.2k
analytics-tracking
Design, audit, and improve analytics tracking systems that produce reliable, decision-ready data.
20 · bundle
flowio
Parse FCS (Flow Cytometry Standard) files v2.0-3.1, extract events as NumPy arrays, read metadata and channels, and convert to CSV or DataFrame for flow cytometry data preprocessing.
30.2k · bundle
analyzing-network-flow-data-with-netflow
Parse NetFlow v9 and IPFIX records to detect volumetric anomalies, port scanning, data exfiltration, and C2 beaconing patterns using the Python netflow library.
24.6k · bundle
whoop
Fetches latest WHOOP recovery, sleep, and strain data and generates daily suggestions.
1 · bundle
youtube-video-api-skill
Extracts structured channel-level and video detail data from a YouTube channel via the BrowserAct API, including metrics like views, likes, comments, and subscriber count.
3.7k · bundle
analyzing-ransomware-network-indicators
Analyze Zeek conn.log and NetFlow data to detect ransomware network indicators including C2 beaconing, TOR exit node connections, data exfiltration, and suspicious DNS patterns.
24.6k · bundle
5-k
Reads and preprocesses 5-minute stock candlestick CSV data, then clusters the time series using tslearn's TimeSeriesKMeans, including data cleaning, percentage change calculation, model training, saving, and representative sample extraction.
559
dropcontact-automation
Automate Dropcontact data enrichment tasks through Composio's Dropcontact toolkit via Rube MCP.
66.9k
jq
Process JSON data from files or standard input using jq filters for extraction, filtering, and transformation.
567 · bundle
polymarket
Queries Polymarket prediction market data via public REST APIs: markets, prices, orderbooks, and history.
2
economic-calendar-fetcher
Fetch upcoming economic events and data releases using the FMP API, including central bank decisions, employment reports, inflation data, GDP releases, and other market-moving indicators for specified date ranges.
2.3k · bundle
polars
High-performance DataFrame library for Python ETL, analytics, and pandas migration. Use for expression-based data manipulation with lazy query optimization, parallel execution, streaming out-of-core processing, Arrow interoperability, and optional GPU execution.
253 · bundle
jq
Query, filter, transform, and aggregate JSON data using jq, with practical patterns for shell pipelines and CLI integration.
253