Bench360 Eval

Evaluates local LLM inference across multiple dimensions, including task-specific quality (e.g., accuracy, F1, ROUGE) and system-level performance (latency, throughput, energy, memory, cold-start) under simulated workloads (single-stream, batch, server). Use when the user wants to benchmark on mmlu, squad_v2, cnn_dailymail, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4bf6567 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bench360-eval commit 4bf65671bf

Frequently asked questions

npx skillmds add qhjqhj00/bench360-eval