Harmbench Asr Eval

Evaluates the robustness of LLM safety defenses against multi-turn human and automated jailbreak attacks. It probes whether current refusal mechanisms and machine unlearning methods can withstand adversarial red teaming aimed at recovering harmful or dual-use knowledge. Use when the user wants to benchmark on HarmBench, WMDP-Bio, or asks about evaluating this task. Reports ASR.

qhjqhj00 7fab1ef 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/harmbench-asr-eval commit 7fab1ef222

Frequently asked questions

npx skillmds add qhjqhj00/harmbench-asr-eval