klue-ner-eval
KLUE: Korean Language Understanding Evaluation — Sungjoon Park et al. (arXiv:2105.09680, 2021)
What this evaluates
Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding.
Datasets
- KLUE-NER — total ?; splits: train (-1), val (-1), test (-1); repo https://github.com/KLUE-benchmark/KLUE
Metrics
F1(primary) — range: [0, 1]- Harmonic mean of precision and recall for entity boundary and label matching.
Input / output format
Input: Korean sentence with tokenized input.
Output: Sequence of BIO/IOB entity tags per token.
Scoring recipe
def compute_f1(predictions, gold):
tp = sum(1 for p, g in zip(predictions, gold) if p == g and p != 'O')
fp = sum(1 for p, g in zip(predictions, gold) if p != g and p != 'O')
fn = sum(1 for p, g in zip(predictions, gold) if p != g and g != 'O')
prec = tp / (tp + fp) if (tp + fp) > 0 else 0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0
return 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
Common pitfalls
- Tokenization mismatches between model BPE and gold character-level annotations.
- Strict boundary matching penalizes minor offset errors heavily.
Evidence (verbatim from paper)
KLUE introduces a comprehensive, ethically designed benchmark for Korean NLU with 8 tasks (Topic Classification, STS, NLI, NER, RE, DP, MRC, DST) built from scratch using diverse, copyright-respected corpora.
Citation
@misc{park2021klue,
title={KLUE: Korean Language Understanding Evaluation},
author={Sungjoon Park et al.},
year={2021},
note={arXiv:2105.09680}
}
- arXiv: 2105.09680