Palm Fewshot Nlp Eval

Evaluates the few-shot and fine-tuned capabilities of large autoregressive language models across a wide range of English NLP benchmarks, including question answering, reading comprehension, common sense reasoning, and natural language inference. It also assesses performance on a large collection of collaborative reasoning and language tasks to probe multi-step reasoning and general language understanding. Use when the user wants to benchmark on English NLP Benchmarks (29 tasks), MMLU, BIG-bench (textual), or asks about evaluating this task. Reports accuracy.

qhjqhj00 edaf2a3 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/palm-fewshot-nlp-eval commit edaf2a3c5c

Frequently asked questions

npx skillmds add qhjqhj00/palm-fewshot-nlp-eval