Model Training & Fine-tuning
-
orchestra-research Bundle Quantizing Models BitsandbytesQuantize LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss using bitsandbytes. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers.
Audited 10.4k -
orchestra-research Bundle Nemo Evaluator SdkEvaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
10.4k -
orchestra-research Bundle Sentence TransformersGenerate high-quality sentence and text embeddings for semantic similarity, clustering, and retrieval using 5000+ pre-trained models. Supports multilingual and domain-specific embeddings for RAG and semantic search.
Audited 10.4k -
orchestra-research Bundle NanogptTrain and experiment with a minimal GPT implementation in ~300 lines of PyTorch, from character-level Shakespeare to GPT-2 scale.
Audited 10.4k -
orchestra-research Bundle Lambda Labs Gpu CloudManage and use Lambda Labs GPU cloud instances for ML training and inference with SSH access, persistent filesystems, and multi-node clusters.
10.4k -
orchestra-research Skill LlamaguardDeploy Meta's LlamaGuard moderation model to filter LLM inputs and outputs across 6 safety categories using HuggingFace, vLLM, or FastAPI.
10.4k -
orchestra-research Bundle Llama CppRun LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.
Audited 10.4k -
orchestra-research Bundle Nemo CuratorGPU-accelerated data curation for LLM training, supporting text, image, video, and audio with fuzzy deduplication, quality filtering, semantic deduplication, PII redaction, and NSFW detection.
Audited 10.4k -
orchestra-research Bundle Optimizing Attention FlashOptimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
Audited 10.4k -
orchestra-research Bundle Fine Tuning With TrlFine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
Audited 10.4k -
orchestra-research Bundle Grpo Rl TrainingExpert guidance for implementing GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.
Audited 10.4k -
orchestra-research Bundle Tensorrt LLMOptimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency on NVIDIA GPUs (A100/H100).
Audited 10.4k -
orchestra-research Bundle Huggingface AccelerateAdd distributed training support to any PyTorch script with minimal code changes using a unified API for DDP, DeepSpeed, FSDP, and mixed precision.
Audited 10.4k -
orchestra-research Bundle Long ContextExtend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques for processing long documents and implementing efficient positional encodings.
Audited 10.4k -
orchestra-research Bundle Moe TrainingTrain Mixture of Experts (MoE) models using DeepSpeed or HuggingFace, covering architectures, routing, load balancing, and expert parallelism.
10.4k -
orchestra-research Bundle Model MergingMerge multiple fine-tuned models using mergekit to combine capabilities without retraining, covering SLERP, TIES-Merging, DARE, Task Arithmetic, linear merging, and production deployment strategies.
10.4k -
orchestra-research Skill Constitutional AITrain AI models to be harmless through self-critique and AI feedback using a set of constitutional principles, without requiring human labels for harmful outputs.
Audited 10.4k -
orchestra-research Bundle Training Llms MegatronTrains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies for maximum GPU efficiency.
10.4k -
orchestra-research Bundle Pytorch Fsdp2Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.
Audited 10.4k -
orchestra-research Bundle Pyvene InterventionsPerform causal interventions on PyTorch models using pyvene's declarative framework for causal tracing, activation patching, and interchange intervention training.
Audited 10.4k -
orchestra-research Bundle Nnsight Remote InterpretabilityRun interpretability experiments on neural network internals using nnsight, with optional NDIF remote execution for massive models.
10.4k -
orchestra-research Bundle Sparse Autoencoder TrainingTrain and analyze Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features for mechanistic interpretability research.
10.4k -
orchestra-research Bundle Speculative DecodingAccelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques for 1.5-3.6× speedup without quality loss.
10.4k -
huggingface Skill Hf Cloud Sagemaker Deployment PlannerPlans and coordinates the deployment of a model to Amazon SageMaker AI, selecting the appropriate pathway (real-time, serverless, async, batch, or Bedrock CMI) based on model type, traffic, latency, and cost constraints.
10.8k -
antigravity Skill Ml EngineerBuild production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks, including model serving, feature engineering, A/B testing, and monitoring.
Audited 42.4k -
google-gemma Bundle Gemma DevSelects the right Gemma model for a task, recommends deployment tooling (Gradio, Transformers.js, Vertex AI, MLX), and applies optimizations like MTP and QAT.
-
systemtce Bundle 03 PerformanceOptimizes Dify workflows and plugins by restructuring graphs, reducing LLM token usage, tuning worker pools, and improving parallel processing.
34 -
majiayu000 Bundle AI Ml TechnologiesCovers AI, machine learning, LLMs, prompt engineering, and blockchain development with code examples and best practices for building AI applications and smart contracts.
567 -
majiayu000 Bundle MlGuides machine learning development with experiment tracking, hyperparameter optimization, model registry, and MLOps pipeline integration.
567 -
majiayu000 Bundle RnaAnnotates single-cell RNA-seq data by scoring marker genes, transferring labels with CellTypist, or reasoning over marker lists with an LLM.
Audited 567 -
majiayu000 Bundle Awq QuantizationQuantize large language models to 4-bit precision using activation-aware weight quantization, reducing memory footprint and speeding up inference with minimal accuracy loss.
Audited 567 -
majiayu000 Bundle DitClassifies HTML pages, forms, and fields using machine learning to detect page types, form types, and field types from HTML content or URLs.
567 -
majiayu000 Bundle Hqq QuantizationQuantize large language models to 8/4/3/2/1-bit precision without calibration data, using multiple optimized backends and integrations with HuggingFace Transformers, vLLM, and PEFT/LoRA.
Audited 567 -
majiayu000 Bundle JaxHigh-performance numerical computing with JAX, covering functional transformations, Flax NNX, and best practices for ML research.
Audited 567 -
majiayu000 Bundle RayScales Python ML workloads across clusters using Ray's distributed tasks, actors, data, and serving capabilities.
Audited 567 -
majiayu000 Bundle Universal Single Cell AnnotatorAnnotates single-cell RNA-seq data by scoring marker genes, transferring labels with CellTypist, or reasoning over cluster markers with an LLM.
Audited 567