Mm Scale Eval

Evaluates vision-language models' ability to perform fine-grained moral reasoning and safety alignment on multimodal scenarios. It probes how well models rank, calibrate, and separate safe from unsafe situations when provided with text, image, or combined modalities. Use when the user wants to benchmark on MM-Scale, or asks about evaluating this task. Reports NDCG@5.

qhjqhj00 529b327 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-scale-eval commit 529b32720c

Frequently asked questions

npx skillmds add qhjqhj00/mm-scale-eval