Medframeqa Eval

This benchmark evaluates multi-image medical visual question answering and clinical reasoning. It probes a model's ability to integrate diagnostic evidence across temporally coherent medical images, detect salient findings, and propagate reasoning chains to answer single-choice questions. Use when the user wants to benchmark on MedFrameQA, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 c180874 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medframeqa-eval commit c18087419c

Frequently asked questions

npx skillmds add qhjqhj00/medframeqa-eval