Temmed Bench Eval

Evaluates large vision-language models' ability to perform temporal reasoning on medical images by analyzing condition changes across multiple clinical visits. It probes capabilities in visual question answering, longitudinal clinical report generation, and selecting relevant image pairs based on temporal context. Use when the user wants to benchmark on TemMed-Bench, or asks about evaluating this task. Reports Avg..

qhjqhj00 e85f8e1 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/temmed-bench-eval commit e85f8e1350

Frequently asked questions

npx skillmds add qhjqhj00/temmed-bench-eval