Head QA V2 Eval

This benchmark evaluates large language models on complex medical reasoning using real Spanish medical licensing exam questions. It probes domain-specific knowledge retention, cross-lingual generalization, and the effectiveness of various inference strategies like prompting, retrieval-augmented generation, and log-probability selection. Use when the user wants to benchmark on HEAD-QA v2, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6ae9a60 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/head-qa-v2-eval commit 6ae9a60d9f

Frequently asked questions

npx skillmds add qhjqhj00/head-qa-v2-eval