Med Tiv Eval

Evaluates a verifier model's ability to distinguish correct from erroneous reasoning traces in medical question-answering tasks. It measures how well tool-integrated reinforcement learning improves factual justification and reduces hallucination compared to static reward models. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, MedXpertQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f0e52a5 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/med-tiv-eval commit f0e52a5552

Frequently asked questions

npx skillmds add qhjqhj00/med-tiv-eval