Medprmbench Eval

This benchmark evaluates the ability of process reward models and general critic models to detect factual, logical, and clinical errors at individual reasoning steps in medical question-answering. It probes step-level correctness verification and case-level chain validation, emphasizing the identification of clinically critical mistakes such as missing contraindications, flawed diagnostic logic, or premature conclusions. Use when the user wants to benchmark on MedPRMBench, or asks about evaluating this task. Reports PRMScore.

qhjqhj00 40423e2 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medprmbench-eval commit 40423e22fb

Frequently asked questions

npx skillmds add qhjqhj00/medprmbench-eval