Instruction Robustness Eval

Evaluates the zero-shot robustness of instruction-tuned language models to variations in instruction phrasing, even when instructions are semantically equivalent. It measures how well models maintain performance on unobserved instruction variants compared to observed ones. Use when the user wants to benchmark on MMLU, BBL, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7a6d965 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/instruction-robustness-eval commit 7a6d965b6c

Frequently asked questions

npx skillmds add qhjqhj00/instruction-robustness-eval