Hrm8k Eval

Evaluates multilingual mathematical reasoning capability, specifically probing whether models can comprehend and solve Korean math problems by leveraging English-as-pivot reasoning to bridge cross-lingual comprehension gaps. Use when the user wants to benchmark on HRM8K, or asks about evaluating this task. Reports pass@1.

qhjqhj00 0381327 2.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hrm8k-eval commit 0381327292

Frequently asked questions

npx skillmds add qhjqhj00/hrm8k-eval