Mti Bench Eval

Evaluates whether large language models can process multiple distinct instructions simultaneously within a single inference call, compared to sequential or batched approaches. It probes reasoning consistency, format adherence, and inference efficiency across a diverse set of 28 NLP tasks. Use when the user wants to benchmark on MTI Bench, or asks about evaluating this task. Reports exact match (EM).

qhjqhj00 33bc7bb 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mti-bench-eval commit 33bc7bb460

Frequently asked questions

npx skillmds add qhjqhj00/mti-bench-eval