shaPO-safety-eval

Evaluates the robustness of LLM safety alignment under in-distribution, cross-domain, and noisy supervision settings. It probes whether geometry-aware optimization preserves safety performance while resisting distribution shift and corrupted preference labels. Use when the user wants to benchmark on PKU-SafeRLHF-30K, HH-RLHF-Safety, Do-Not-Answer, HarmBench, SaladBench, or asks about evaluating this task. Reports Win Rate (WR).

qhjqhj00 b7e97df 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/shaPO-safety-eval commit b7e97dffa5

Frequently asked questions

npx skillmds add qhjqhj00/shapo-safety-eval