# Klue Ner Eval

> Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding. Use when the user wants to benchmark on KLUE-NER, or asks about evaluating this task. Reports F1.

- Skill: `qhjqhj00/klue-ner-eval` (Agent Skill)
- Install (CLI): `npx skillmds add qhjqhj00/klue-ner-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/klue-ner-eval/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/qhjqhj00/klue-ner-eval

---


# klue-ner-eval

> KLUE: Korean Language Understanding Evaluation — Sungjoon Park et al. (arXiv:2105.09680, 2021)

## What this evaluates

Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding.

## Datasets

- **KLUE-NER** — total ?; splits: train (-1), val (-1), test (-1); repo https://github.com/KLUE-benchmark/KLUE

## Metrics

- `F1` **(primary)** — range: [0, 1]
  - Harmonic mean of precision and recall for entity boundary and label matching.

## Input / output format

**Input**: Korean sentence with tokenized input.

**Output**: Sequence of BIO/IOB entity tags per token.

## Scoring recipe

```python
def compute_f1(predictions, gold):
    tp = sum(1 for p, g in zip(predictions, gold) if p == g and p != 'O')
    fp = sum(1 for p, g in zip(predictions, gold) if p != g and p != 'O')
    fn = sum(1 for p, g in zip(predictions, gold) if p != g and g != 'O')
    prec = tp / (tp + fp) if (tp + fp) > 0 else 0
    rec = tp / (tp + fn) if (tp + fn) > 0 else 0
    return 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
```

## Common pitfalls

- Tokenization mismatches between model BPE and gold character-level annotations.
- Strict boundary matching penalizes minor offset errors heavily.

## Evidence (verbatim from paper)

> KLUE introduces a comprehensive, ethically designed benchmark for Korean NLU with 8 tasks (Topic Classification, STS, NLI, NER, RE, DP, MRC, DST) built from scratch using diverse, copyright-respected corpora.

## Citation

```bibtex
@misc{park2021klue,
  title={KLUE: Korean Language Understanding Evaluation},
  author={Sungjoon Park et al.},
  year={2021},
  note={arXiv:2105.09680}
}
```

- arXiv: 2105.09680

