Meissa Medical Eval

Evaluates a 4B multi-modal medical agentic model's ability to perform clinical reasoning, tool use, and multi-step interaction across radiology, pathology, and clinical domains. Probes strategy selection (when to use tools vs direct reasoning) and execution policy under various agent frameworks. Use when the user wants to benchmark on MIMIC-CXR-VQA, ChestAgentBench, PathVQA, SLAKE, VQA-RAD, OmniMed, MedXpertQA, MedQA, PubMedQA, NEJM, NEJM Ext., MIMIC-IV, MedQA Ext., or asks about evaluating this task. Reports accuracy.

qhjqhj00 0fcc655 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/meissa-medical-eval commit 0fcc65592e

Frequently asked questions

npx skillmds add qhjqhj00/meissa-medical-eval