Msr Align Safety Eval

Evaluates whether fine-tuning vision-language models on policy-grounded safety reasoning improves their ability to refuse unsafe multimodal prompts while preserving general multimodal reasoning capabilities. Use when the user wants to benchmark on BeaverTails-V, MM-SafetyBench, SPA-VL Eval, MME-CoT, MM-Vet, or asks about evaluating this task. Reports safety rate.

qhjqhj00 8549afb 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/msr-align-safety-eval commit 8549afbaa5

Frequently asked questions

npx skillmds add qhjqhj00/msr-align-safety-eval