Safeclassip Toxic Defense Eval

Evaluates the ability of Large Vision-Language Models (LVLMs) to detect and refuse harmful visual content without modifying the base model architecture. It measures both safety defense effectiveness on toxic inputs and the preservation of utility on benign inputs. Use when the user wants to benchmark on Toxic Image Categories (Porn, Bloody, Insulting, Alcohol, Cigarette, Gun, Knife, Neutral), or asks about evaluating this task. Reports DSR.

qhjqhj00 f3487fc 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/safeclassip-toxic-defense-eval commit f3487fce5b

Frequently asked questions

npx skillmds add qhjqhj00/safeclassip-toxic-defense-eval