Dialectalarabicmmlu Eval

Evaluates large language models' ability to understand and reason across multiple Arabic dialects and standard Arabic across diverse academic and professional domains. It measures dialectal generalization and sensitivity to linguistic context by comparing performance under default, dialect-conditioned, and dialect-identification prompts. Use when the user wants to benchmark on DialectalArabicMMLU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 adef64d 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dialectalarabicmmlu-eval commit adef64d4f1

Frequently asked questions

npx skillmds add qhjqhj00/dialectalarabicmmlu-eval