ETL & Pipelines

153 skills
gabrielmoreira
seq-wrangler
Runs NGS read QC, alignment, and BAM processing, wrapping FastQC, BWA/Bowtie2/Minimap2, SAMtools, and MultiQC for automated read-to-BAM workflows.
17 · bundle
gabrielmoreira
ncbi-datasets
Downloads genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
17 · bundle
auto-skiller
data-scraping
Builds a configurable scraping agent that collects data from APIs, HTML, or RSS, enriches it with Gemini AI scoring, and stores results in Notion, Google Sheets, Supabase, or local files.
1 · bundle
diegosouzapw
dbt
Provides dbt patterns for data transformation and analytics engineering, including model structures, incremental models, and testing.
54 · bundle
oyi77
dbt-transform
Transforms raw data into analytics-ready models using dbt, covering models, tests, macros, sources, snapshots, documentation, and packages.
10
phoroth
jq
Query, filter, transform, and aggregate JSON data using jq, with practical patterns for shell pipelines and CLI integration.
3
lucaspmarie-a11y
polars
Process in-memory datasets with Polars' expression API, lazy evaluation, and parallel execution, including pandas migration patterns and I/O for CSV, Parquet, and JSON.
5
qhjqhj00
ray-data
Process large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
3 · bundle
iamanacarolinarezende
database
Guides database design, implementation, query optimization, migrations, data pipeline development, and operations across SQL and NoSQL platforms.
0