Mm Safetybench Eval

Evaluates the safety and detoxification capabilities of multimodal large language models by measuring the fraction of harmful responses across various toxicity categories, while also assessing continuous toxicity severity and general multimodal reasoning capability. Use when the user wants to benchmark on MM-SafetyBench, or asks about evaluating this task. Reports Harmful Rate (HaR).

qhjqhj00 9acfd55 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-safetybench-eval commit 9acfd55968

Frequently asked questions

npx skillmds add qhjqhj00/mm-safetybench-eval