Scam Typographic Robustness Eval

This benchmark evaluates the typographic robustness of vision-language models (VLMs) and large vision-language models (LVLMs) by measuring their susceptibility to adversarial handwritten or synthetic text inserted into images. It probes whether models can correctly identify the primary object in an image despite the presence of misleading attack words, revealing vulnerabilities in multimodal alignment and text-visual reasoning. Use when the user wants to benchmark on SCAM, or asks about evaluating this task. Reports accuracy.

qhjqhj00 3b7f660 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scam-typographic-robustness-eval commit 3b7f66008e

Frequently asked questions

npx skillmds add qhjqhj00/scam-typographic-robustness-eval