Omh Inference Serving

[omh] OMH Inference Serving workflow: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says: inference-serving, inference serving, serve this model, serve the model, model serving, serving endpoint, vllm, llama.cpp.

rlaope Updated

File contents

rlaope/oh-my-hermes/tree/main/skills/omh-inference-serving commit 7941ba22b1

Frequently asked questions

npx skillmds@latest add rlaope/omh-inference-serving-3