Omh Inference Serving

[omh] OMH Inference Serving workflow: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says: inference-serving, inference serving, serve this model, serve the model, model serving, serving endpoint, vllm, llama.cpp.

rlaope 1286cb9 3 files · 13.8 KB Updated

File contents

rlaope/oh-my-hermes commit 1286cb9cca

Frequently asked questions

npx skillmds@latest add rlaope/omh-inference-serving