Wizardlm Eval

Probes instruction-following capability on complex, real-world prompts across diverse domains like coding, math, reasoning, and formatting. It measures how well models handle demanding, multi-step tasks compared to baselines through blind pairwise human comparison. Use when the user wants to benchmark on WizardEval, or asks about evaluating this task. Reports win_rate.

qhjqhj00 b394c4b 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wizardlm-eval commit b394c4b543

Frequently asked questions

npx skillmds add qhjqhj00/wizardlm-eval