Gke Inference

>- Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for generic batch jobs or HPC task queues (use gke-batch-hpc instead).

thedixitjain 586b5a8 7.6 KB Updated 2 repo stars

File contents

thedixitjain/the-mega-skill-library/tree/main/library/data-science-and-ml/gke-inference commit 586b5a8799

Frequently asked questions

npx skillmds add thedixitjain/gke-inference