Gradsafe Jailbreak Detection Eval

Evaluates the ability to detect jailbreak or unsafe prompts in LLM inputs using gradient-based analysis. It probes zero-shot and adapted detection capabilities against established moderation APIs and LLM-based detectors. Use when the user wants to benchmark on ToxicChat, XSTest, or asks about evaluating this task. Reports AUPRC.

qhjqhj00 791de46 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gradsafe-jailbreak-detection-eval commit 791de46140

Frequently asked questions

npx skillmds add qhjqhj00/gradsafe-jailbreak-detection-eval