Mobile Mmlu Eval

Evaluates language models' understanding of mobile-specific domains and tasks under on-device constraints. It probes the models' ability to answer multiple-choice questions across 80 real-world mobile domains, emphasizing practical usability, privacy, and personalization in daily mobile interactions. Use when the user wants to benchmark on Mobile-MMLU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 ab0fb13 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mobile-mmlu-eval commit ab0fb13080

Frequently asked questions

npx skillmds add qhjqhj00/mobile-mmlu-eval