Multimodal Red Teaming Eval

Evaluates the safety and harm susceptibility of multimodal large language models (MLLMs) when exposed to adversarial prompts across different input modalities (text-only vs. image-text). It measures how effectively these prompts bypass safety filters and the severity of the resulting harmful outputs. Use when the user wants to benchmark on Multimodal Adversarial Benchmark, or asks about evaluating this task. Reports Attack Success Rate (ASR).

qhjqhj00 d0e7e71 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-red-teaming-eval commit d0e7e7172e

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-red-teaming-eval