Bionpars Bench Eval

Evaluates a Persian biomedical large language model's ability to generate accurate, domain-specific long-form answers and summaries. It probes subject-specific knowledge acquisition, knowledge synthesis, and evidence-based reasoning by comparing model outputs against human-written biomedical references. Use when the user wants to benchmark on BioPars-BENCH, or asks about evaluating this task. Reports BERTScore.

qhjqhj00 7ec499f 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bionpars-bench-eval commit 7ec499f2b3

Frequently asked questions

npx skillmds add qhjqhj00/bionpars-bench-eval