Results for “gpu-cluster”
35 skillsairunway-aks-setup
Walks users from a bare AKS cluster to a running AI model deployment, covering cluster verification, controller install, GPU assessment, provider setup, and first deployment.
2.7k · bundle
tao-run-on-slurm
Submit and manage TAO training, evaluation, and inference jobs on SLURM GPU clusters over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed storage.
2.2k · bundle
gke-cluster-creation
Creates GKE clusters with golden path Autopilot defaults, supporting Standard and GPU workloads. Guides through project, region, and networking inputs, then provisions and verifies cluster settings.
14.4k
lambda-labs-gpu-cloud
Manage and use Lambda Labs GPU cloud instances for ML training and inference with SSH access, persistent filesystems, and multi-node clusters.
10.4k · bundle
More results
lambda-labs-gpu-cloud
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
1 · bundle
tao-setup-nvidia-gpu-host
Checks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
2.2k · bundle
gke-upgrades
Plans, executes, and validates Google Kubernetes Engine (GKE) cluster upgrades and maintenance operations for both Standard and Autopilot clusters, producing upgrade plans, checklists, and runbooks with gcloud commands.
14.4k · bundle
gke-batch-hpc
Runs batch processing and high-performance computing (HPC) workloads on Google Kubernetes Engine (GKE), including job queues, parallel processing, and MPI workloads.
14.4k
gke-basics
Routes to specialized GKE sub-skills for cluster management, networking, security, scaling, and more on Google Kubernetes Engine.
14.4k · bundle
gcp-gke
Manages Google Kubernetes Engine clusters via gcloud CLI, covering cluster discovery, node pool management, workload analysis, autopilot configuration, and upgrade planning.
7
lambda-labs-gpu-cloud
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
0 · bundle
gke-cluster-config
Configure gke cluster config operations. Auto-activating skill for GCP Skills. Triggers on: gke cluster config, gke cluster config Part of the GCP Skills skill category. Use when configuring systems or services. Trigger with phrases like "gke cluster config", "gke config", "gke".
4
gke-storage
**Trigger**: Use when working with GKE Storage — Google Kubernetes Engine configuration and management.
1
gke-basics
Set up and manage Google Kubernetes Engine clusters, node pools, workloads, networking, and storage with Autopilot defaults.
1
gke-inference
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
14.4k
gke-cluster-autoscaler
Provides guidance on enabling and optimizing GKE Cluster Autoscaler, including Node Auto Provisioning, troubleshooting scale-up/down issues, and best practices for capacity management.
14.4k · bundle
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses for Spot VMs with on-demand fallback, GPU/TPU targeting, machine family selection, and zone colocation.
14.4k · bundle
gke-security
Hardens Google Kubernetes Engine (GKE) clusters with Workload Identity, Secret Manager, RBAC, Binary Authorization, Network Policies, and Pod Security Standards.
14.4k · bundle
gke-storage
Configures GKE storage including PVCs, PersistentVolumes, Filestore, and GCS FUSE with best practices for production workloads.
14.4k
tao-run-on-kubernetes
Submits TAO container jobs as single-pod Kubernetes Jobs with NVIDIA GPU scheduling on EKS, GKE, AKS, or on-prem clusters.
2.2k · bundle
gke-multitenancy
Plans and configures multi-tenancy on GKE, covering namespace isolation, RBAC planning, resource quotas, LimitRanges, network isolation, and cost allocation.
14.4k
mcore-run-on-slurm
Launch distributed Megatron-LM training jobs on a SLURM cluster with a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules, container conventions, monitoring, and per-rank failure diagnosis.
2.2k · bundle
gke-backup-dr
Protects stateful GKE workloads by configuring backup plans, restore workflows, and disaster recovery using Backup for GKE.
14.4k
gke-networking
Plans, configures, and manages GKE networking including private clusters, VPC-native configurations, Gateway API, DNS, ingress/egress, Dataplane V2, and IP planning.
14.4k
gke-reliability
Configures GKE workload reliability with PodDisruptionBudgets, health probes, topology spread constraints, and graceful shutdown patterns.
14.4k
memorystore-config
Configure memorystore config operations. Auto-activating skill for GCP Skills. Triggers on: memorystore config, memorystore config Part of the GCP Skills skill category. Use when configuring systems or services. Trigger with phrases like "memorystore config", "memorystore config", "memorystore".
4
gcp-alloydb
Manages Google AlloyDB for PostgreSQL clusters and instances via gcloud CLI, covering discovery, instance analysis, query insights, backup configuration, and maintenance windows.
7
gcp-cloud-run
Specialized skill for building production-ready serverless applications on GCP. Covers Cloud Run services (containerized), Cloud Run Functions (event-driven), cold start optimization, and event-driven architecture with Pub/Sub.
505 · bundle
gcp-specialist-ia
Expert en infrastructure GCP (GKE, Cloud Run, BigQuery, Pub/Sub, Firestore, Cloud Functions)
6
cloud-gcp
Operates Google Cloud Platform resources via the gcloud CLI, covering compute instances, storage buckets, BigQuery, Cloud Run, GKE, IAM, and billing. It checks the environment, authenticates, confirms destructive operations, and diagnoses common errors.
2
kubernetes
Kubernetes
128 · bundle
google-cloud
Deploy, monitor, and manage GCP services with battle-tested patterns.
12
emulating-cloud-attacks-with-stratus-red-team
Detonate granular AWS, Azure, GCP, and Kubernetes attack techniques to validate detections with Stratus Red Team.
24.6k · bundle
gcp
Deploys and manages Google Cloud Platform backends for Flutter apps, covering GKE clusters, Cloud Run services, Cloud SQL databases, Cloud Storage buckets, and Firebase integration.
4
launch-nemo-rl
Launch, monitor, stop, and debug NeMo-RL recipes on a Kubernetes cluster using the nrl-k8s CLI, supporting ephemeral and long-lived RayCluster modes.
2.2k · bundle