Hardbench Eval

Evaluates whether LLMs are vulnerable to draft-based co-authoring jailbreaks that exploit collaborative writing contexts to elicit harmful completions. It probes the model's ability to detect concealed malicious intent in incomplete drafts and assesses the trade-off between safety refusal and writing utility. Use when the user wants to benchmark on HarDBench, or asks about evaluating this task. Reports Harmfulness Score (HS).

qhjqhj00 609e17d 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hardbench-eval commit 609e17db0f

Frequently asked questions

npx skillmds add qhjqhj00/hardbench-eval