Evaluating Kubernetes Performance Genai

Design and optimize Kubernetes-native GenAI inference platforms using Kueue job queuing, Dynamic Accelerator Slicer (DAS) GPU partitioning, and Gateway API Inference Extension (GAIE) with llm-d for multi-stage AI pipelines. Use when: 'set up Kubernetes for AI inference', 'configure Kueue for batch GPU jobs', 'partition GPUs with MIG slicing on Kubernetes', 'optimize LLM inference routing on Kubernetes', 'build a Whisper-to-LLM pipeline on K8s', 'reduce TTFT latency for LLM serving'.

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/evaluating-kubernetes-performance-genai commit 1589036c80

Frequently asked questions

npx skillmds@latest add ndpvt-web/evaluating-kubernetes-performance-genai