Universal Ner V2 Eval

Evaluates multilingual named entity recognition (NER) capabilities across 22 languages and 30 datasets, probing both in-language performance and cross-lingual transfer. It also benchmarks large language models as annotators against human inter-annotator agreement to assess guideline adherence and annotation quality. Use when the user wants to benchmark on UNER v2, or asks about evaluating this task. Reports micro F1.

qhjqhj00 c2411c5 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/universal-ner-v2-eval commit c2411c5fc4

Frequently asked questions

npx skillmds add qhjqhj00/universal-ner-v2-eval