Smmile Eval

Evaluates the ability of multimodal large language models (MLLMs) to perform in-context learning (ICL) in medical domains. It probes how effectively models leverage provided image-question-answer demonstrations to answer new clinical queries, while also measuring robustness to irrelevant examples, recency bias, and the gap between automated and expert clinical judgment. Use when the user wants to benchmark on SMMILE, SMMILE++, or asks about evaluating this task. Reports LLM-as-a-Judge.

qhjqhj00 242c866 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/smmile-eval commit 242c86611b

Frequently asked questions

npx skillmds add qhjqhj00/smmile-eval