Wildjailbreak Eval

Evaluates the safety and robustness of language models against adversarial jailbreak attacks. It probes whether models can correctly refuse harmful requests while avoiding over-refusal on benign prompts, specifically under stealthy, adversarially composed prompts. Use when the user wants to benchmark on WILDJAILBREAK, or asks about evaluating this task. Reports Attack success rate (ASR).

qhjqhj00 2e7e08b 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wildjailbreak-eval commit 2e7e08b705

Frequently asked questions

npx skillmds add qhjqhj00/wildjailbreak-eval