Omni Safetybench Eval

Evaluates the safety alignment and refusal capabilities of Audio-Visual Large Language Models (OLLMs) when exposed to harmful unimodal, dual-modal, and omni-modal inputs. It specifically probes whether models maintain consistent safety boundaries across modality combinations and reveals vulnerabilities in cross-modal comprehension-aware safety. Use when the user wants to benchmark on Omni-SafetyBench, or asks about evaluating this task. Reports Safety-score.

qhjqhj00 05443d7 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omni-safetybench-eval commit 05443d7ab2

Frequently asked questions

npx skillmds add qhjqhj00/omni-safetybench-eval