Super Naturalinstructions Eval

Evaluates instruction-following and cross-task generalization capabilities of language models on a massive, diverse benchmark of 1,616 NLP tasks spanning 76 task types and 55 languages. It measures how well models trained on a mix of tasks perform on unseen tasks when given natural language instructions. Use when the user wants to benchmark on Super-NaturalInstructions, or asks about evaluating this task. Reports human evaluation metric.

qhjqhj00 3ba1996 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/super-naturalinstructions-eval commit 3ba1996b17

Frequently asked questions

npx skillmds add qhjqhj00/super-naturalinstructions-eval