Jailbreakbench Eval

Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).

qhjqhj00 7dd69e1 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/jailbreakbench-eval commit 7dd69e1670

Frequently asked questions

npx skillmds add qhjqhj00/jailbreakbench-eval