Opt Iml Bench Eval

Evaluates instruction-tuned language models on generalization across 1,991 NLP tasks spanning 100+ categories. It probes zero-shot and few-shot (5-shot) performance on held-out categories, unseen tasks within seen categories, and fully supervised tasks, measuring both generation quality and classification accuracy. Use when the user wants to benchmark on OPT-IML Bench, or asks about evaluating this task. Reports Rouge-L.

qhjqhj00 a53ba4a 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/opt-iml-bench-eval commit a53ba4a5cb

Frequently asked questions

npx skillmds add qhjqhj00/opt-iml-bench-eval