Fakebench Eval

Evaluates large multimodal models on explainable fake image detection across closed-ended classification and open-ended reasoning tasks. It probes the models' ability to accurately classify image authenticity and generate evidence-based, interpretable justifications using visual and textual forensic cues. Use when the user wants to benchmark on FakeBench, or asks about evaluating this task. Reports Accuracy (ACC).

qhjqhj00 13dd721 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/fakebench-eval commit 13dd7212db

Frequently asked questions

npx skillmds add qhjqhj00/fakebench-eval