Arahahealthqa Eval

This benchmark evaluates Arabic language models on healthcare-related question answering, specifically probing their ability to classify mental health conditions and generate culturally appropriate medical advice. It tests both discriminative capabilities (multi-label classification and multiple-choice selection) and generative capabilities (open-ended response generation) in clinical and mental health contexts. Use when the user wants to benchmark on AraHealthQA, or asks about evaluating this task. Reports Weighted-F1.

qhjqhj00 efd5a5a 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/arahahealthqa-eval commit efd5a5ab02

Frequently asked questions

npx skillmds add qhjqhj00/arahahealthqa-eval