ETL & Pipelines
153 skillsseq-wrangler
Runs NGS read QC, alignment, and BAM processing, wrapping FastQC, BWA/Bowtie2/Minimap2, SAMtools, and MultiQC for automated read-to-BAM workflows.
17 · bundle
ncbi-datasets
Downloads genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
17 · bundle
data-scraping
Builds a configurable scraping agent that collects data from APIs, HTML, or RSS, enriches it with Gemini AI scoring, and stores results in Notion, Google Sheets, Supabase, or local files.
1 · bundle
dbt
Provides dbt patterns for data transformation and analytics engineering, including model structures, incremental models, and testing.
54 · bundle
dbt-transform
Transforms raw data into analytics-ready models using dbt, covering models, tests, macros, sources, snapshots, documentation, and packages.
10
jq
Query, filter, transform, and aggregate JSON data using jq, with practical patterns for shell pipelines and CLI integration.
3
polars
Process in-memory datasets with Polars' expression API, lazy evaluation, and parallel execution, including pandas migration patterns and I/O for CSV, Parquet, and JSON.
5
ray-data
Process large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
3 · bundle
database
Guides database design, implementation, query optimization, migrations, data pipeline development, and operations across SQL and NoSQL platforms.
0