Odd Eval

This benchmark evaluates a model's ability to perform multi-label classification on clinical electronic health record (EHR) notes to detect nine categories of Opioid-Related Aberrant Behaviors (ORABs). It probes the model's capacity to identify both confirmed and suggested aberrant behaviors, as well as auxiliary opioid-related signals, under conditions of significant label imbalance. Use when the user wants to benchmark on ODD, or asks about evaluating this task. Reports macro average AUPRC.

qhjqhj00 f0654aa 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/odd-eval commit f0654aa0a0

Frequently asked questions

npx skillmds add qhjqhj00/odd-eval