Mm Safetybench++ Eval

Evaluates contextual safety in multi-modal large language models by measuring how well they refuse harmful queries while correctly answering safe ones, with a focus on whether their safety reasoning aligns with the given context. It probes the model's ability to avoid over-defensive refusals on benign inputs while maintaining high response quality. Use when the user wants to benchmark on MM-SafetyBench++, or asks about evaluating this task. Reports Contextual Correctness Rate (CCR).

qhjqhj00 6b42cb4 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-safetybench++-eval commit 6b42cb48d5

Frequently asked questions

npx skillmds add qhjqhj00/mm-safetybench-eval-2