Egida Safety Eval

Evaluates the robustness of LLMs against jailbreaking attacks after safety alignment. It measures how well models refuse harmful prompts across diverse topics and attack styles, while also tracking unintended side effects like over-refusal and general capability degradation. Use when the user wants to benchmark on Egida, or asks about evaluating this task. Reports ASR.

qhjqhj00 4d8d11f 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/egida-safety-eval commit 4d8d11f972

Frequently asked questions

npx skillmds add qhjqhj00/egida-safety-eval