Wildguard Moderation Eval

This evaluation protocol assesses the safety moderation capabilities of LLMs and dedicated moderation models. It probes their ability to detect harmful content in user prompts, classify harmful or safe model responses, and identify whether a model appropriately refuses unsafe requests across multiple risk categories. Use when the user wants to benchmark on ToxicChat, OpenAI Mod, AegisSafetyTest, SimpleSafetyTests, Harmbench Prompt, Harmbench Resp, BeaverTails, SafeRLHF, XSTest-Resp, WildGuardTest, or asks about evaluating this task. Reports F1 score.

qhjqhj00 3519b48 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wildguard-moderation-eval commit 3519b48015

Frequently asked questions

npx skillmds add qhjqhj00/wildguard-moderation-eval