Bj Benchmark Eval

This benchmark evaluates vision-language and large language models on clinical reasoning for musculoskeletal disorders. It probes capabilities ranging from medical knowledge recall and unimodal interpretation to open-ended multimodal diagnosis, treatment planning, and text-image inconsistency detection. The protocol highlights the performance gap between structured multiple-choice questions and complex, free-form clinical reasoning tasks. Use when the user wants to benchmark on B&J benchmark, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 0196d9d 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bj-benchmark-eval commit 0196d9d5af

Frequently asked questions

npx skillmds add qhjqhj00/bj-benchmark-eval