Vllm Serving Setup

Design, deploy, and tune vLLM v0.18.2 inference serving on EKS with PagedAttention v2, Multi-LoRA, FP8 KV Cache, Chunked Prefill, and Continuous Batching. Produces Helm values.yaml, PodMonitor, HPA, and kubectl validation steps for production agentic workloads.

aws-samples f79a9f6 4.4 KB Updated

File contents

aws-samples/sample-oh-my-aidlcops/tree/main/plugins/ai-infra/skills/vllm-serving-setup commit f79a9f6ca1

Frequently asked questions

npx skillmds@latest add aws-samples/vllm-serving-setup