LLM Serving Eval

Evaluates the throughput, latency, and scalability of LLM inference serving systems under varying request rates and context lengths. It probes how efficiently a system manages KV cache, batching, and resource allocation for both short and long-context instruction-following workloads. Use when the user wants to benchmark on Alpaca, LongBench, or asks about evaluating this task. Reports Throughput.

qhjqhj00 1b99113 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llm-serving-eval commit 1b9911339d

Frequently asked questions

npx skillmds add qhjqhj00/llm-serving-eval