Plugins

3 plugins

Results for “data-pipeline”

26 skills
More results
phoroth
jq
Query, filter, transform, and aggregate JSON data using jq, with practical patterns for shell pipelines and CLI integration.
3
oyi77
prefect-flows
Orchestrates Python data pipelines with Prefect flows, tasks, retries, caching, parallel execution, and deployments to work pools.
10
nvidia
data-designer
Build synthetic datasets and data generation pipelines using the Data Designer library.
2.2k · bundle
nvidia
nemo-data-designer-plugin
Build synthetic datasets and data generation pipelines using the Data Designer library.
2.2k · bundle
redpanda-data
connect
Build streaming data pipelines with Redpanda Connect using declarative YAML configs, Bloblang mappings, and component discovery. Covers running, linting, and dry-running pipelines.
6 · bundle
leandrobenjaminl
etl-pipelines
Construye pipelines ETL/ELT con Pandas: extracción, transformación y carga de datos con logging, manejo de errores, idempotencia y opciones de orquestación.
0 · bundle
jeffallan
ml-pipeline
Designs and implements production-grade ML pipeline infrastructure: configures experiment tracking, creates orchestration DAGs, builds feature store schemas, deploys model registries, and automates retraining and validation workflows.
10.4k · bundle
gabrielmoreira
flow-bio
Authenticate, browse pipelines, samples, and projects, upload data, launch pipeline executions, and check run status on any Flow.bio instance via CLI.
17 · bundle
github
bigquery-pipeline-audit
Audits Python + BigQuery pipelines for cost safety, idempotency, and production readiness, returning a structured report with exact patch locations.
36.2k
phoroth
database
Guides database design, implementation, optimization, migration, pipeline development, quality, and operations across SQL and NoSQL platforms.
3
leandrobenjaminl
shared-git-data
Sets up Git-based version control for data science projects, handling notebooks, datasets, and pipelines with DVC and nbstripout.
0
ssrjkk
dbt-etl
Guides the setup and use of dbt for extract-transform-load workflows.
2 · bundle
jeffallan
spark-engineer
Write, optimize, and debug Apache Spark jobs for high-performance distributed data processing, ETL pipelines, and big data workloads.
10.4k · bundle
redpanda-data
connect-debugging
Diagnoses and validates Redpanda Connect pipelines using linting, dry-run connection tests, logging, metrics, tracing, and health endpoints, including enterprise feature troubleshooting.
6 · bundle
antigravity
polars
Provides a fast in-memory DataFrame library for datasets that fit in RAM, with lazy evaluation, parallel execution, and an Apache Arrow backend for ETL pipelines and analytics.
42.4k
muratcankoylan
book-sft-pipeline
Convert books into supervised fine-tuning datasets and train style-transfer models that replicate an author's voice.
16.9k · bundle
qhjqhj00
ray-data
Process large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
3 · bundle
k-dense-ai
latchbio-integration
Build and deploy bioinformatics workflows as serverless pipelines on the Latch platform using Python decorators, cloud data management, and GPU support.
30.2k · bundle
k-dense-ai
nextflow
Build, run, and debug Nextflow data pipelines and nf-core workflows end to end, covering processes, channels, operators, configuration, testing, and deployment to HPC or cloud.
30.2k · bundle
k-dense-ai
dnanexus-integration
Build and deploy apps/applets on the DNAnexus cloud genomics platform, manage data objects, run workflows, and use the dxpy Python SDK for genomics pipeline development and execution.
30.2k · bundle