Continual Safety Alignment Eval

Evaluates a model's ability to maintain safety alignment and task performance during sequential continual fine-tuning across multiple domains. It probes whether gradient-based sample selection prevents safety degradation (elastic reversion) and catastrophic forgetting while preserving general capabilities. Use when the user wants to benchmark on AdvBench, HarmBench, TruthfulQA, ARC-C, BoolQ, HellaSwag, Winogrande, GSM8K, MedMCQA, Squad_v2, or asks about evaluating this task. Reports ASR.

qhjqhj00 601a1bd 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/continual-safety-alignment-eval commit 601a1bd0f7

Frequently asked questions

npx skillmds add qhjqhj00/continual-safety-alignment-eval