Blue Eval

Evaluates the cross-domain generalization and transfer learning capabilities of pre-trained language models across ten diverse biomedical and clinical NLP tasks. It probes sentence similarity, named entity recognition, relation extraction, document classification, and natural language inference to measure how well domain-specific pre-training captures clinical and biomedical semantics. Use when the user wants to benchmark on MedSTS, BIOSSES, BC5CDR-disease, BC5CDR-chemical, ShARe/CLEFE, DDI, ChemProt, i2b2 2010, HoC, MedNLI, or asks about evaluating this task. Reports Total Score (Macro-average).

qhjqhj00 487c230 4.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/blue-eval commit 487c230ccb

Frequently asked questions

npx skillmds add qhjqhj00/blue-eval