Ervqa Eval

This benchmark evaluates the clinical readiness of Large Vision Language Models (LVLMs) for emergency room monitoring tasks. It probes their ability to generate accurate, clinically cautious, and semantically entailed long-form answers from medical images, while identifying specific failure modes like hallucinations and overconfidence. Use when the user wants to benchmark on ERVQA, or asks about evaluating this task. Reports Entailment Score.

qhjqhj00 1aa01b6 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ervqa-eval commit 1aa01b6dba

Frequently asked questions

npx skillmds add qhjqhj00/ervqa-eval