Vl Rewardbench Eval

Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.

qhjqhj00 a6fb43c 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vl-rewardbench-eval commit a6fb43c361

Frequently asked questions

npx skillmds add qhjqhj00/vl-rewardbench-eval