Adversarial Nibbler Eval

This benchmark evaluates the robustness of text-to-image models against implicitly adversarial prompts—subtle, non-obvious text inputs that bypass automated safety filters to generate harmful images. It probes the gap between human safety perception and machine safety classification, highlighting long-tail failure modes and context-dependent vulnerabilities in generative AI. Use when the user wants to benchmark on Nibbler, or asks about evaluating this task. Reports false_negative_rate.

qhjqhj00 d3bea41 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/adversarial-nibbler-eval commit d3bea41646

Frequently asked questions

npx skillmds add qhjqhj00/adversarial-nibbler-eval