First Inference Eval

Evaluates the performance and scalability of a federated inference scheduling framework under varying request loads, measuring throughput, latency, and auto-scaling capabilities on distributed HPC resources. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports Request throughput (req/s).

qhjqhj00 ac408aa 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/first-inference-eval commit ac408aa20e

Frequently asked questions

npx skillmds add qhjqhj00/first-inference-eval