Pmc Mi Bench Eval

Evaluates multi-modal large language models on medical reasoning tasks involving compound figures, single images, and text-only prompts. It probes the model's ability to synthesize cross-modal information, perform clinical diagnosis, and generate accurate medical explanations across diverse imaging modalities and specialties. Use when the user wants to benchmark on PMC-MI-Bench, or asks about evaluating this task. Reports BLEU@4, Accuracy.

qhjqhj00 19bddee 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pmc-mi-bench-eval commit 19bddeeb8e

Frequently asked questions

npx skillmds add qhjqhj00/pmc-mi-bench-eval