Jailbreak Attack Eval

This protocol evaluates the robustness of large language models against automated jailbreak attacks. It measures how effectively generated or human-crafted prompts can bypass safety filters to elicit prohibited or harmful responses. Use when the user wants to benchmark on 100 questions from two open datasets [6,37], or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 0341f40 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/jailbreak-attack-eval commit 0341f4038c

Frequently asked questions

npx skillmds add qhjqhj00/jailbreak-attack-eval