Metaicl Fewshot Eval

Evaluates the few-shot adaptation capability of language models on a diverse set of NLP tasks. It measures how well a model generalizes to unseen tasks after being fine-tuned on automatically extracted few-shot examples from web tables. Use when the user wants to benchmark on Min et al. (2021) Tasks, CROSSFIT, UNIFIEDQA, or asks about evaluating this task. Reports mean Dev Tasks score.

qhjqhj00 59930b0 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/metaicl-fewshot-eval commit 59930b0921

Frequently asked questions

npx skillmds add qhjqhj00/metaicl-fewshot-eval