Guardrail Robustness Eval

Evaluates the robustness and generalization of LLM safety guardrails against adversarial jailbreak prompts, measuring their ability to correctly classify harmful vs. benign inputs under both known benchmark distributions and novel, contextually framed attacks. Use when the user wants to benchmark on Adversarial Guardrail Benchmark, or asks about evaluating this task. Reports Overall Accuracy.

qhjqhj00 ae8bf0d 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/guardrail-robustness-eval commit ae8bf0d8ef

Frequently asked questions

npx skillmds add qhjqhj00/guardrail-robustness-eval