Detailverifybench Eval

Evaluates multimodal large language models' ability to pinpoint erroneous content at the token level within long-form image captions. It probes whether models can distinguish between visually grounded facts and hallucinated details by localizing specific tokens that contradict the input image. Use when the user wants to benchmark on DetailVerifyBench, or asks about evaluating this task. Reports token-level F1.

qhjqhj00 6f2cf8b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/detailverifybench-eval commit 6f2cf8b673

Frequently asked questions

npx skillmds add qhjqhj00/detailverifybench-eval