Medarabiq Eval

Evaluates large language models on Arabic medical reasoning and dialogue across multiple-choice, fill-in-the-blank, and open-ended Q&A tasks. It probes factual accuracy, domain-specific knowledge, and robustness to linguistic variations and injected biases in healthcare contexts. Use when the user wants to benchmark on MedArabiQ, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 4306277 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medarabiq-eval commit 43062779ef

Frequently asked questions

npx skillmds add qhjqhj00/medarabiq-eval