ETL & Pipelines

153 skills
github
sql-server-table-reconciliation
Compare identical tables across two SQL Server instances using Python with mssql-python and Apache Arrow, detecting missing rows, column mismatches, schema drift, and generating a reconciliation report.
36.2k · bundle
github
dataverse-python-advanced-patterns
Generate production-ready Python code for Dataverse SDK with advanced patterns including error handling, batch operations, OData optimization, and Pandas integration.
36.2k
jeffallan
spark-engineer
Write, optimize, and debug Apache Spark jobs for high-performance distributed data processing, ETL pipelines, and big data workloads.
10.4k · bundle
tradermonty
skill-idea-miner
Extract, score, and backlog skill idea candidates from Claude Code session logs for a weekly skill generation pipeline.
2.3k · bundle
tradermonty
edge-candidate-agent
Convert daily market observations into structured research tickets and export validated candidate specs for a trading strategy pipeline.
2.3k · bundle
tradermonty
edge-concept-synthesizer
Clusters raw detection tickets into reusable edge concepts with thesis, invalidation signals, and strategy playbooks before strategy design.
2.3k · bundle
redpanda-data
connect-debugging
Diagnoses and validates Redpanda Connect pipelines using linting, dry-run connection tests, logging, metrics, tracing, and health endpoints, including enterprise feature troubleshooting.
6 · bundle
redpanda-data
connect-cdc-salesforce
Streams Salesforce change data capture and platform events into Redpanda or Kafka using Redpanda Connect's salesforce_cdc input, covering setup, configuration, and operational details.
6 · bundle
pranavnagrecha
fhir-data-mapping
Maps FHIR R4 clinical resources (Patient, Observation, Condition, CarePlan, CodeableConcept) to Salesforce Health Cloud objects, including prerequisite configuration and cardinality handling.
15 · bundle
leandrobenjaminl
db-admin
Administra bases de datos PostgreSQL, MySQL, Redis y SQLite: diseña esquemas, optimiza queries, configura migraciones, replicación y backups.
0
leandrobenjaminl
shared-git-data
Sets up Git-based version control for data science projects, handling notebooks, datasets, and pipelines with DVC and nbstripout.
0
nimoqup046-collab
database
Guides database design, implementation, optimization, migration, pipeline development, and operations across SQL and NoSQL platforms.
2
zero-yx
music-ingestion-publisher
Ingests music into a LanceDB via ncmdump-rs and sf-cli, searching and downloading from Netease Cloud Music or Bilibili, decrypting local NCM files, and extracting metadata, lyrics, and cover art.
0 · bundle
gabrielmoreira
wren
Provides a discovery stub for the Wren CLI, a semantic SQL layer over 22+ databases, with commands to install, set up, connect data sources, generate MDL projects, enrich context, and deploy GenBI dashboards.
17
gabrielmoreira
wgs-prs
Takes raw whole-genome sequencing FASTQ files or a pre-existing VCF through variant calling, quality control, and polygenic risk score computation using the PGS Catalog.
17 · bundle
gabrielmoreira
flow-bio
Authenticate, browse pipelines, samples, and projects, upload data, launch pipeline executions, and check run status on any Flow.bio instance via CLI.
17 · bundle
gabrielmoreira
galaxy-bridge
Discovers and executes bioinformatics tools from the Galaxy ecosystem via natural language, with multi-signal scoring, workflow templates, and reproducibility bundles.
17 · bundle
auto-skiller
pyragify
Converts code repositories and document directories into semantically-chunked text files optimized for NotebookLM ingestion, with support for config files and incremental processing.
1 · bundle
oyi77
prefect-flows
Orchestrates Python data pipelines with Prefect flows, tasks, retries, caching, parallel execution, and deployments to work pools.
10
tools-only
153-dxpy-bae649e0
Provides Python bindings to interact with the DNAnexus platform, enabling file uploads, job management, and API calls.
7 · bundle
schattenspiegel
pyarrow-python
Write, review, debug, test, or optimize Python code using PyArrow arrays, schemas, tables, compute kernels, datasets, Parquet, and Arrow IPC.
0 · bundle
nvidia
tao-generate-image-grounding
Generates phrase-grounded bounding box annotations from image-caption pairs using a VLM, producing cleaned captions, referring expressions, and pixel-space bounding boxes.
2.2k · bundle
nvidia
digital-health-clinical-asr-build
Curates clinical-specialty term lists, generates IPA-tagged synthetic audio via TTS, and produces NeMo-format manifests for ASR benchmark evaluation.
2.2k · bundle
adobe
page-import
Import a single webpage from any URL into canonical EDS block format — structured HTML that authors edit in DA. Scrapes the page, analyzes structure, maps to existing blocks, and generates HTML for immediate local preview.
142 · bundle
agentskillexchange
dbt-mcp-server
Guides installation of dbt Core or dbt Cloud CLI and points to official documentation for getting started with dbt.
28
redpanda-data
connect-cdc-mysql
Streams MySQL or MariaDB row-level changes into Redpanda or Kafka via the mysql_cdc input, covering binlog setup, snapshots, checkpointing, per-table routing, AWS RDS IAM auth, and Enterprise lakehouse destinations.
6 · bundle
redpanda-data
sql-federated-queries
Query external data from Oxla — Kafka topics via catalogs, Apache Iceberg tables, and S3/GCS/Azure parquet/ORC files — alongside native Oxla tables. Use when querying Kafka topics with CREATE KAFKA CATALOG or CREATE REDPANDA CATALOG, reading Apache Iceberg tables with the catalog=>path.table syntax, loading or.
6 · bundle
pranavnagrecha
bulk-api-patterns
Implements Bulk API 2.0 REST calls for ingest and query jobs, covering CSV format requirements, locator pagination, and v1 vs v2 selection.
15 · bundle
pranavnagrecha
salesforce-data
Routes to the right Salesforce data skill package for data model, migration, bulk loads, query optimization, deduplication, and archival tasks.
15 · bundle
pranavnagrecha
gift-history-import
Migrates historical donation or gift records into Salesforce NPSP using the NPSP Data Importer (BDI), covering DataImport__c staging, payment mapping, soft credits via Opportunity Contact Roles, GAU allocation, and campaign attribution.
15 · bundle
leandrobenjaminl
data-archive
Documenta, versiona y cierra proyectos de análisis de datos para que queden ordenados y reproducibles en el futuro.
0
scoheart
firecrawl-scrape
Extracts clean, LLM-optimized markdown from any URL, including JavaScript-rendered SPAs, with support for concurrent scraping of multiple URLs and options like main-content-only extraction and custom output formats.
2
joshuashepherd
markitdown
Converts files and office documents to Markdown, supporting PDF, DOCX, PPTX, XLSX, images with OCR, audio with transcription, HTML, CSV, JSON, XML, ZIP, YouTube URLs, and EPubs.
1 · bundle
luokai0
a-stock-get
Collects Chinese A-share stock data, fetching stock lists and daily, weekly, and monthly K-line history into a local SQLite database for quantitative analysis.
10 · bundle
jorcan
calc
Creates, edits, converts, and automates LibreOffice Calc spreadsheets in ODS format, including formulas, charts, and batch processing.
0 · bundle
jorcan
database
Guides database design, implementation, optimization, migration, pipeline development, quality, and operations across SQL and NoSQL platforms.
0 · bundle