# Twnertc Eval

> Evaluates the annotation quality of the TWNERTC corpus for Turkish named entity recognition (NER) and text categorization (TC) by comparing automated labels against human-verified ground truths. It measures how well coarse-grained and fine-grained entity types, as well as domain categories, align with human judgment across different noise-reduction post-processing variants. Use when the user wants to benchmark on TWNERTC, or asks about evaluating this task. Reports F1-Score.

- Skill: `qhjqhj00/twnertc-eval` (Agent Skill)
- Install (CLI): `npx skillmds add qhjqhj00/twnertc-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/twnertc-eval/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Productivity
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/qhjqhj00/twnertc-eval

---


# twnertc-eval

> Automatically Annotated Turkish Corpus for Named Entity Recognition and Text Categorization using Large-Scale Gazetteers — H. Bahadir Sahin et al. (2017) (arXiv:1702.02363, 2017)

## What this evaluates

Evaluates the annotation quality of the TWNERTC corpus for Turkish named entity recognition (NER) and text categorization (TC) by comparing automated labels against human-verified ground truths. It measures how well coarse-grained and fine-grained entity types, as well as domain categories, align with human judgment across different noise-reduction post-processing variants.

## Datasets

- **TWNERTC** — total ?; splits: test_ner (10000), test_tc (2000)

## Metrics

- `F1-Score` **(primary)** — range: [0, 1]
  - Harmonic mean of precision and recall, calculated per entity type (Person, Org, Loc, Misc) and averaged across types. Formula: F1 = 2 * (Precision * Recall) / (Precision + Recall).
- `Top-K Agreement` — range: [0, 1]
  - Percentage of instances where the human-annotated ground truth label appears within the top K predictions provided by the automated system or model.

## Input / output format

**Input**: Sentences containing automatically annotated entities (with coarse- or fine-grained types) or domain labels.

**Output**: Human-annotated ground truth labels: for NER, exact entity type tags (e.g., person, organization, location, misc, O); for TC, a ranked list of up to 5 domain categories.

## Scoring recipe

```python
# Coarse-Grained NER: Precision, Recall, F1
preds = [p for p, g in zip(predictions, golds) if g != 'O']
golds = [g for p, g in zip(predictions, golds) if g != 'O']
tp = sum(1 for p, g in zip(preds, golds) if p == g)
fp = sum(1 for p, g in zip(preds, golds) if p != g)
fn = sum(1 for p, g in zip(preds, golds) if g != p)
prec = tp / (tp + fp) if (tp + fp) > 0 else 0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0
f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0

# Fine-Grained NER & TC: Top-K Agreement
def top_k_agreement(preds, golds, k):
    matches = sum(1 for p, g in zip(preds, golds) if g in p[:k])
    return matches / len(preds)
```

## Common pitfalls

- Fine-grained NER ground truth is derived from an English model (FIGER) adapted to Turkish, which may introduce language-specific biases not present in pure human annotation.
- The 'Misc' entity type consistently yields the lowest F1-Score due to its broad coverage and high susceptibility to mismatching, which is expected but often overlooked.
- Ground truth for fine-grained and TC tasks is constructed via human ranking of model-generated candidates rather than manual annotation from scratch, making the evaluation a measure of candidate quality rather than absolute labeling accuracy.

## Evidence (verbatim from paper)

> We present precision, recall and F-score values in Table[4]... In Table[5], we present the F1-scores of trained models on each test set... Table[6] presents the human annotators evaluation on automated fine-grained NER datasets, given FIGER predictions as possible ground-truths. Annotators rank the provided ground-truths and we check their ranking agreements. Eventually, top-1 agreement is hard to fulfill since our gazetteers contains thousands of entity types, and an entity may have more than ten different possible options. On the other hand, top-5 agreements provide promising results considering the amount of possible ground-truths.

## Citation

```bibtex
@misc{sahin2017twnertc,
  title={Automatically Annotated Turkish Corpus for Named Entity Recognition and Text Categorization using Large-Scale Gazetteers},
  author={H. Bahadir Sahin et al. (2017)},
  year={2017},
  note={arXiv:1702.02363}
}
```

- arXiv: 1702.02363

