Medeval Eval

Evaluates language models on multi-level (sentence/document) and multi-task (NLU/NLG) medical benchmarks across diverse clinical domains. It probes a model's ability to perform clinical text classification, report code prediction, and medical report summarization using both fine-tuned PLMs and prompted LLMs. Use when the user wants to benchmark on MedEval, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6cc9923 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medeval-eval commit 6cc9923531

Frequently asked questions

npx skillmds add qhjqhj00/medeval-eval