Plugins
1 pluginResults for “datasets”
18 skillsNcbi Datasets
Downloads genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.
17 · bundle
Shared Git Data
Sets up Git-based version control for data science projects, handling notebooks, datasets, and pipelines with DVC and nbstripout.
0
Data Designer
Build synthetic datasets and data generation pipelines using the Data Designer library.
2.2k · bundle
Nemo Data Designer Plugin
Build synthetic datasets and data generation pipelines using the Data Designer library.
2.2k · bundle
Tao Convert Dataset Format
Converts NVIDIA TAO DAFT datasets between supported formats using the `tao-daft convert` CLI.
2.2k · bundle
Dask
Scale pandas and NumPy workflows to larger-than-memory datasets using parallel and distributed computing.
30.2k · bundle
More results
Apify Actor Runner
Runs Apify cloud actors for structured web scraping and exports datasets to S3, with input schema validation and webhook notifications.
28
Pyarrow Python
Write, review, debug, test, or optimize Python code using PyArrow arrays, schemas, tables, compute kernels, datasets, Parquet, and Arrow IPC.
0 · bundle
Cupynumeric Parallel Data Load
Load sharded datasets (npy, Parquet, HDF5, raw binary) into distributed cuPyNumeric arrays using manual partitioning and Legate task launches.
2.2k · bundle
Book Sft Pipeline
Convert books into supervised fine-tuning datasets and train style-transfer models that replicate an author's voice.
16.9k · bundle
Big Data
Designs and implements big data architectures, processes large-scale datasets with distributed systems, and optimizes data pipelines for throughput using Hadoop, Spark, and cloud platforms.
1
Pysam
Read, write, and analyze genomic datasets including SAM/BAM/CRAM alignments, VCF/BCF variants, and FASTA/FASTQ sequences using a Pythonic interface to htslib.
253 · bundle
Lamindb
Manages biological datasets and models with LaminDB, covering setup, artifact registration, querying, lineage tracking, validation, ontology annotation, collections, branches, storage, and workflow integrations.
253 · bundle
Ray Data
Process large-scale ML datasets with distributed streaming execution across CPU/GPU, supporting Parquet, CSV, JSON, images, and integration with PyTorch, TensorFlow, and Ray Train.
10.4k · bundle
Polars
Process in-memory datasets with Polars' expression API, lazy evaluation, and parallel execution, including pandas migration patterns and I/O for CSV, Parquet, and JSON.
5
Polars
Provides a fast in-memory DataFrame library for datasets that fit in RAM, with lazy evaluation, parallel execution, and an Apache Arrow backend for ETL pipelines and analytics.
42.4k
Vaex
Process and analyze large tabular datasets (billions of rows) that exceed available RAM using lazy, out-of-core DataFrames with fast aggregations, visualization, and machine learning integration.
30.2k · bundle
Ray Data
Process large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
3 · bundle