Crossguard Multimodal Safety Eval

Evaluates the robustness of multimodal LLMs against explicit and implicit jailbreak attacks while measuring their utility on benign queries. It probes whether a defense model can successfully refuse harmful image-text prompts without over-restricting safe inputs. Use when the user wants to benchmark on JailBreakV, VLGuard, FigStep, MM-SafetyBench, SIUO, MMBench, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 d7d836d 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/crossguard-multimodal-safety-eval commit d7d836d278

Frequently asked questions

npx skillmds add qhjqhj00/crossguard-multimodal-safety-eval