Massive Eval

Evaluates multilingual natural language understanding capabilities, specifically intent classification and slot filling, across 51 typologically diverse languages. It measures model robustness to different scripts, spacing conventions, and zero-shot cross-lingual transfer scenarios. Use when the user wants to benchmark on MASSIVE, or asks about evaluating this task. Reports exact match accuracy.

qhjqhj00 5a58da7 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/massive-eval commit 5a58da72d0

Frequently asked questions

npx skillmds add qhjqhj00/massive-eval