Safer Safety Eval

Evaluates LLM safety alignment and robustness against adversarial jailbreak attacks, particularly in scientific domains. It also measures the model's ability to maintain general helpfulness, truthfulness, and avoid over-refusal on benign queries. Use when the user wants to benchmark on AdvBench, HarmBench, StrongReject, SciKnowEval (L4), SciSafeEval, LabSafety Bench (Hard), GSM8K, MT-Bench, MMLU, GPQA, SimpleQA, XsTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 90d56d9 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/safer-safety-eval commit 90d56d97ae

Frequently asked questions

npx skillmds add qhjqhj00/safer-safety-eval