Refinedweb Zero Shot Eval

Evaluates the zero-shot generalization capability of autoregressive language models across multiple task aggregates. It measures how well models trained on raw web data perform on downstream tasks without any fine-tuning or prompt engineering. Use when the user wants to benchmark on Eleuther AI LM evaluation harness (zero-shot aggregates), or asks about evaluating this task. Reports zero-shot accuracy.

qhjqhj00 9674f33 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/refinedweb-zero-shot-eval commit 9674f3347c

Frequently asked questions

npx skillmds add qhjqhj00/refinedweb-zero-shot-eval