Packs

2 packs

Results for “kubernetes”

43 skills
More results
nvidia
dynamo-recipe-runner
Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes for model serving with GPU support.
2.2k · bundle
trailofbits
debug-buttercup
Diagnose pod crashes, restart loops, Redis failures, resource pressure, and other service misbehavior in the crs namespace on Kubernetes.
6k · bundle
google
gke-golden-path
Provides GKE golden path configuration defaults, production readiness checklists, and cluster default patterns for designing and verifying GKE clusters.
14.4k · bundle
google
gke-security
Hardens Google Kubernetes Engine (GKE) clusters with Workload Identity, Secret Manager, RBAC, Binary Authorization, Network Policies, and Pod Security Standards.
14.4k · bundle
oracle
oke-gva-deployer
Deploy and configure Generic VNIC Attachment (GVA) for OCI Kubernetes Engine (OKE), including node pool creation with secondary VNIC profiles and workload mapping.
736 · bundle
google
gke-batch-hpc
Runs batch processing and high-performance computing (HPC) workloads on Google Kubernetes Engine (GKE), including job queues, parallel processing, and MPI workloads.
14.4k
nvidia
launch-nemo-rl
Launch, monitor, stop, and debug NeMo-RL recipes on a Kubernetes cluster using the nrl-k8s CLI, supporting ephemeral and long-lived RayCluster modes.
2.2k · bundle
google
gke-upgrades
Plans, executes, and validates Google Kubernetes Engine (GKE) cluster upgrades and maintenance operations for both Standard and Autopilot clusters, producing upgrade plans, checklists, and runbooks with gcloud commands.
14.4k · bundle
google
gke-cluster-creation
Creates GKE clusters with golden path Autopilot defaults, supporting Standard and GPU workloads. Guides through project, region, and networking inputs, then provisions and verifies cluster settings.
14.4k
nvidia
tao-run-platform
Submit and monitor GPU training jobs on Brev, SLURM, Docker, or Kubernetes using the TAO Execution SDK, with job handles, S3 I/O wrapping, and multi-node distributed training.
2.2k · bundle
nvidia
tao-setup-nvidia-gpu-host
Checks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
2.2k · bundle
microsoft
azure-cloud-migrate
Assess and migrate cross-cloud workloads to Azure with reports and code conversion. Supports Lambda to Functions, Beanstalk/Heroku/App Engine to App Service, Fargate/Kubernetes/Cloud Run/Spring Boot to Container Apps.
2.7k · bundle
google
gke-inference
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
14.4k
nvidia
rag-blueprint
Deploy, configure, troubleshoot, and manage NVIDIA RAG Blueprint deployments across Docker, Helm, and library setups.
2.2k · bundle
google
gke-backup-dr
Protects stateful GKE workloads by configuring backup plans, restore workflows, and disaster recovery using Backup for GKE.
14.4k
google
gke-storage
Configures GKE storage including PVCs, PersistentVolumes, Filestore, and GCS FUSE with best practices for production workloads.
14.4k
google
gke-multitenancy
Plans and configures multi-tenancy on GKE, covering namespace isolation, RBAC planning, resource quotas, LimitRanges, network isolation, and cost allocation.
14.4k
google
gke-reliability
Configures GKE workload reliability with PodDisruptionBudgets, health probes, topology spread constraints, and graceful shutdown patterns.
14.4k
github
qdrant-monitoring-setup
Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.
36.2k
google
gke-cost
Optimize GKE costs by rightsizing workloads, configuring Spot VMs, selecting machine types, and using Committed Use Discounts.
14.4k
nvidia
aiq-deploy
Installs, deploys, runs, validates, troubleshoots, and stops NVIDIA AI-Q Blueprint infrastructure for local or self-hosted servers.
2.2k · bundle
dotnet
dump-collect
Configure and collect crash dumps for modern .NET applications (CoreCLR and NativeAOT) on Linux, macOS, and Windows, including containers.
4k · bundle
google
gke-scaling
Configures GKE autoscaling with HPA, VPA, and Node Auto-Provisioning using golden path defaults for cost optimization.
14.4k · bundle
google
gke-observability
Configures GKE observability with Cloud Logging, Cloud Monitoring, and managed Prometheus for monitoring, logging, and metrics collection.
14.4k
nvidia
dynamo-troubleshoot
Diagnose failed or unhealthy Dynamo deployments by collecting a read-only debug bundle, classifying failures, and providing step-by-step remediation guidance.
2.2k · bundle
google
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses for Spot VMs with on-demand fallback, GPU/TPU targeting, machine family selection, and zone colocation.
14.4k · bundle
nvidia
tao-run-inference-service
Start, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
2.2k · bundle
nvidia
dynamo-router-starter
Start or patch Dynamo router modes and run router endpoint smoke checks for round-robin, KV-aware, least-loaded, or device-aware routing.
2.2k · bundle
oracle
oke-multihome-deployer
Deploy and verify Multus-based pod multihoming on OKE clusters with existing GVA secondary VNIC node pools, including discovery, manifest generation, and connectivity validation.
736 · bundle
github
devops-rollout-plan
Generate comprehensive rollout plans with preflight checks, step-by-step deployment, verification signals, rollback procedures, and communication plans for infrastructure and application changes.
36.2k