Perceptionprocessbench Eval

Evaluates a vision-language process reward model's ability to detect step-level visual grounding and reasoning errors in structured multimodal reasoning traces. The benchmark specifically probes whether the model can distinguish between correct and subtly mutated perception steps that are designed to be challenging for automated error detection. Use when the user wants to benchmark on PerceptionProcessBench, or asks about evaluating this task. Reports step-level correctness.

qhjqhj00 01c5660 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/perceptionprocessbench-eval commit 01c56607d7

Frequently asked questions

npx skillmds add qhjqhj00/perceptionprocessbench-eval