Model Perf Binary Search

Find the maximum sustainable QPS of an LLM inference service that meets a p50 e2e latency SLO using online_replay.py and a binary search. Use for maximum-QPS benchmarks, SLO-based performance tests, and optional serving configuration or feature tuning on local OpenAI-compatible servers.

Saddss Updated

File contents

Saddss/cursor-skills/tree/main/skills/model-perf-binary-search commit 8865d2909f

Frequently asked questions

npx skillmds@latest add saddss/model-perf-binary-search