Mart Safety Eval

Evaluates an LLM's ability to refuse harmful or unsafe requests while maintaining helpfulness on benign prompts. It probes safety alignment through automatic reward-model scoring and human flagging of violations across in-distribution and out-of-domain benchmarks. Use when the user wants to benchmark on SafeEval, HelpEval, AlpacaEval, Anthropic Harmless, or asks about evaluating this task. Reports violation_rate.

qhjqhj00 659b298 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mart-safety-eval commit 659b298862

Frequently asked questions

npx skillmds add qhjqhj00/mart-safety-eval