Medical Summarization Eval

This evaluation probes the ability of large language models to generate accurate and faithful summaries of medical texts under high out-of-vocabulary (OOV) conditions. It specifically measures how tokenization fragmentation and domain-specific terminology affect summarization quality and concept preservation across multiple medical benchmarks. Use when the user wants to benchmark on PubMedQA, EBM, BioASQ-M, BioASQ-S, or asks about evaluating this task. Reports Rouge-L.

qhjqhj00 36ef6f5 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-summarization-eval commit 36ef6f58f8

Frequently asked questions

npx skillmds add qhjqhj00/medical-summarization-eval