Gptaraeval Eval

Evaluates large language models on Arabic natural language understanding and generation across 44 tasks and over 60 datasets, covering both Modern Standard Arabic and dialectal varieties. Use when the user wants to benchmark on GPTAraEval Benchmark Suite, or asks about evaluating this task. Reports macro-F1.

qhjqhj00 ce124c9 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/gptaraeval-eval commit ce124c9de8

Frequently asked questions

npx skillmds add qhjqhj00/gptaraeval-eval