P3 Niv2 Eval

This evaluation protocol assesses the zero-shot generalization capability of instruction-tuned language models across diverse NLP tasks. It measures how well a model trained on a selected subset of instruction-tuning datasets performs on held-out tasks from the same meta-datasets and external benchmarks, focusing on both classification accuracy and text generation quality. Use when the user wants to benchmark on P3 (Public Pool of Prompts), NIV2 (SuperNaturalInstructions V2), Big-Bench, Big-Bench Hard (BBH), or asks about evaluating this task. Reports ACC.

qhjqhj00 0495db0 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/p3-niv2-eval commit 0495db0701

Frequently asked questions

npx skillmds add qhjqhj00/p3-niv2-eval