Courtguard Safety Eval

Evaluates LLM safety guardrails and policy-adaptation frameworks on their ability to correctly identify harmful, toxic, or policy-violating content across diverse attack vectors. It probes robustness against automated jailbreaks, over-refusal in benign contexts, and zero-shot adaptability to out-of-domain policy enforcement. Use when the user wants to benchmark on AdvBenchM, WildGuard, HarmBench, JailJudge, PKU-SafeRLHF, ToxicChat, BeaverTails, XSTest, PAN Wikipedia Vandalism Corpus 2010, Human-Verified Attack Suite Dataset, or asks about evaluating this task. Reports Accuracy, F1-Score (macro-averaged).

qhjqhj00 debb2f8 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/courtguard-safety-eval commit debb2f86c1

Frequently asked questions

npx skillmds add qhjqhj00/courtguard-safety-eval