Sparrow Alignment Eval

Evaluates the alignment, factual grounding, and rule-following capabilities of dialogue agents through human preference comparisons. It also measures resilience to adversarial probing for specific harm rules and the quality of evidence-supported responses. Use when the user wants to benchmark on ELI5 + Free Dialogue Test Set, or asks about evaluating this task. Reports Three-model preference rate.

qhjqhj00 b2db344 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sparrow-alignment-eval commit b2db344cf2

Frequently asked questions

npx skillmds add qhjqhj00/sparrow-alignment-eval