Dove Eval

This evaluation probes the robustness and prompt sensitivity of large language models on multiple-choice benchmarks by measuring how performance varies across hundreds of millions of intent-preserving prompt perturbations across multiple dimensions. Use when the user wants to benchmark on DOVE, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 fecfb8b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dove-eval commit fecfb8b1a7

Frequently asked questions

npx skillmds add qhjqhj00/dove-eval