Instruction Tuning Eval

Evaluates the instruction-following capability and alignment (helpfulness, honesty, harmlessness) of instruction-tuned LLMs on unseen tasks across English and Chinese. Use when the user wants to benchmark on User-Oriented-Instructions-252, Vicuna-Instructions-80, Unnatural Instructions, or asks about evaluating this task. Reports Relative Score (GPT-4).

qhjqhj00 dcd054e 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/instruction-tuning-eval commit dcd054e1c3

Frequently asked questions

npx skillmds add qhjqhj00/instruction-tuning-eval