Few Shot Nlp Eval

Evaluates the zero-shot and few-shot capabilities of large language models across a diverse suite of NLP, reasoning, and commonsense benchmarks. It measures how efficiently a model scales with compute and whether additional training objectives unlock emergent reasoning abilities. Use when the user wants to benchmark on GPT-3 suite, BigBench Emergent Suite, Commonsense QA benchmarks, Closed-book QA benchmarks, or asks about evaluating this task. Reports average score.

qhjqhj00 7c87a3a 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/few-shot-nlp-eval commit 7c87a3a847

Frequently asked questions

npx skillmds add qhjqhj00/few-shot-nlp-eval