Safeanchor Safety Eval

This evaluation probes a model's ability to retain safety alignment and refusal capabilities while undergoing sequential continual domain adaptation across medical, legal, and coding tasks. It measures cumulative safety erosion and domain performance retention compared to unconstrained fine-tuning baselines. Use when the user wants to benchmark on HarmBench, TruthfulQA, BBQ, WildGuard, MedQA, LegalBench, CodeAlpaca, HumanEval, MMLU, or asks about evaluating this task. Reports Safety Score.

qhjqhj00 017d59e 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/safeanchor-safety-eval commit 017d59e7db

Frequently asked questions

npx skillmds add qhjqhj00/safeanchor-safety-eval