Multimodal Reward Benchmarks Eval

Evaluates multimodal reward models on preference ranking tasks across image and video domains, measuring how well they score or rank candidate responses compared to ground-truth preferences. It compares multi-response scoring against single-response baselines and generative judges, while also assessing inference efficiency and downstream policy optimization stability. Use when the user wants to benchmark on VL-RewardBench, Multimodal RewardBench, MM-RLHF RewardBench, MR2Bench-Image, VideoRewardBench, MR2Bench-Video, or asks about evaluating this task. Reports pairwise accuracy.

qhjqhj00 c9d83c7 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-reward-benchmarks-eval commit c9d83c7b55

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-reward-benchmarks-eval