Hed Benchmark Eval

This benchmark evaluates whether large language models and automated essay scoring systems can correctly distinguish between harmful essays (containing toxic or discriminatory content) and argumentative essays (which present controversial but non-harmful viewpoints). It also measures safety alignment by tracking refusal rates and the tendency to redirect harmful prompts into ethical, argumentative responses. Use when the user wants to benchmark on HED benchmark, or asks about evaluating this task. Reports POR.

qhjqhj00 8d0f5d0 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hed-benchmark-eval commit 8d0f5d0b4b

Frequently asked questions

npx skillmds add qhjqhj00/hed-benchmark-eval