Safebench Asr Eval

Evaluates the vulnerability of aligned multimodal LLMs to universal adversarial image attacks that bypass safety filters. It measures how often a single optimized image forces the model to generate unsafe or affirmative responses across diverse text prompts. Use when the user wants to benchmark on SafeBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 badbfa2 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/safebench-asr-eval commit badbfa2e66

Frequently asked questions

npx skillmds add qhjqhj00/safebench-asr-eval