Kokushimd 10 Eval

This benchmark evaluates large language models' ability to reason through Japanese national healthcare licensing examinations across ten medical professions. It probes domain-specific clinical knowledge, multimodal image interpretation, and high-stakes decision-making under strict, profession-specific passing criteria. Use when the user wants to benchmark on KokushiMD-10, or asks about evaluating this task. Reports accuracy.

qhjqhj00 81c64b0 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kokushimd-10-eval commit 81c64b037c

Frequently asked questions

npx skillmds add qhjqhj00/kokushimd-10-eval