Mt Bench Eval

Evaluates the conversational quality and instruction-following capability of aligned language models across multiple knowledge domains. It also measures whether alignment fine-tuning causes regression in base reasoning, truthfulness, and commonsense capabilities. Use when the user wants to benchmark on MT-Bench, Open LLM Leaderboard Benchmarks, or asks about evaluating this task. Reports MT-Bench Average Score.

qhjqhj00 c8695d5 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mt-bench-eval commit c8695d5e32

Frequently asked questions

npx skillmds add qhjqhj00/mt-bench-eval