Caremedeval Eval

This benchmark evaluates large language models' ability to perform critical appraisal and methodological reasoning on biomedical scientific articles. It probes whether models can correctly identify study design flaws, statistical limitations, and biases by answering multiple-choice questions derived from authentic French medical exams. The evaluation specifically measures both exact correctness and partial reasoning accuracy under varying context conditions. Use when the user wants to benchmark on CareMedEval, or asks about evaluating this task. Reports Exact Match Ratio (EMR).

qhjqhj00 71e151c 4.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/caremedeval-eval commit 71e151c98e

Frequently asked questions

npx skillmds add qhjqhj00/caremedeval-eval