Zero Shot Generalization Eval

Evaluates a model's ability to generalize to unseen natural language tasks without task-specific fine-tuning or prompt tuning. It probes zero-shot performance across traditional NLP benchmarks and novel BIG-bench tasks using accuracy. Use when the user wants to benchmark on BIG-bench & Held-out NLP Tasks, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2a7e5a1 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/zero-shot-generalization-eval commit 2a7e5a1938

Frequently asked questions

npx skillmds add qhjqhj00/zero-shot-generalization-eval