Tmmluplus Eval

Evaluates foundation models' multitask language understanding and reasoning capabilities in Traditional Chinese across diverse academic subjects including STEM, social sciences, humanities, and other domains. Use when the user wants to benchmark on TMMLU+, or asks about evaluating this task. Reports average accuracy (%).

qhjqhj00 224ddf0 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tmmluplus-eval commit 224ddf0299

Frequently asked questions

npx skillmds add qhjqhj00/tmmluplus-eval