Sea Helm Eval

Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.

qhjqhj00 dff5963 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sea-helm-eval commit dff5963dec

Frequently asked questions

npx skillmds add qhjqhj00/sea-helm-eval