Benchmark Tune

Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations. Trigger for requests mentioning benchmark tune, tuning tok/s, ctx_size tradeoffs, mmap or mlock tuning, speculative decoding, MTP, ngram, draft models, or replacing old gpu tune usage.

mesh-llm 155ee59 2 files · 6.6 KB Updated

File contents

mesh-llm/mesh-llm/tree/main/.agents/skills/benchmark-tune commit 155ee59546

Frequently asked questions

npx skillmds@latest add mesh-llm/benchmark-tune