Ares Safety Eval

This evaluation protocol assesses the safety alignment and general capability of large language models after undergoing an adaptive red-teaming and repair process. It probes the model's ability to refuse harmful or unsafe prompts while maintaining performance on standard knowledge and reasoning benchmarks, and measures the false refusal rate to ensure utility is preserved. Use when the user wants to benchmark on RedTeam, StrongReject, HarmBench, PKU-SafeRLHF, XSTest, MMLU, GSM8K, TruthfulQA, AlpacaEval, RewardBench, or asks about evaluating this task. Reports Safety Rate.

qhjqhj00 bba0963 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ares-safety-eval commit bba09631c5

Frequently asked questions

npx skillmds add qhjqhj00/ares-safety-eval