Wingpt 3.0 Benchmark Eval

Evaluates large language models on comprehensive medical reasoning, clinical calculation, and general cognitive capabilities. It probes domain-specific knowledge application, diagnostic reasoning, and complex problem-solving in real-world clinical and academic settings. Use when the user wants to benchmark on MedCalc, MedReMCQ, CMMLU, MATH-500, MedQA-USMLE, MedMCQA, PubMedQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f11741f 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wingpt-3.0-benchmark-eval commit f11741f639

Frequently asked questions

npx skillmds add qhjqhj00/wingpt-3-0-benchmark-eval