finnish-nlp-eval
Towards Fully Bilingual Deep Language Modeling — Chang et al. (2020) (arXiv:2010.11639, 2020)
What this evaluates
Evaluates Finnish NLP capabilities across four core tasks: part-of-speech tagging, named entity recognition, dependency parsing, and text classification on domain-specific and out-of-domain corpora. It probes a model's ability to handle morphologically rich language and varying text registers.
Datasets
- Finnish NLP Benchmarks — total ?; splits: train (-1), dev (-1), test (-1)
Metrics
UPOS — range: [0, 1]
- Unambiguous Part-Of-Speech tag accuracy for POS tagging on Universal Dependencies treebanks.
F-score — range: [0, 1]
- Mention-level F-score for Named Entity Recognition using IOB annotations.
LAS — range: [0, 1]
- Labeled Attachment Score for dependency parsing (accuracy of predicted heads and relations).
Accuracy — range: [0, 1]
- Proportion of correctly classified documents for text classification.
Average (primary) — range: [0, 1]
- Algebraic mean of the four task scores (POS, NER, parsing, classification).
Input / output format
Input: Tokenized Finnish text, typically truncated to at most 256 tokens.
Output: Task-specific predictions: POS tags, IOB entity tags, dependency heads/relations, or document class labels.
Scoring recipe
pos_score = upos_accuracy(y_true, y_pred)
ner_score = mention_level_f1(y_true, y_pred)
parse_score = labeled_attachment_score(y_true, y_pred)
cls_score = document_accuracy(y_true, y_pred)
average = (pos_score + ner_score + parse_score + cls_score) / 4
Common pitfalls
- Uses gold segmentation for dependency parsing instead of predicted segmentation.
- Truncates all documents to 256 tokens, which may disadvantage models with longer context windows.
- Averages different metrics (UPOS, F1, LAS, Accuracy) without normalization, which can skew the final score.
Evidence (verbatim from paper)
The metric for POS tagging is UPOS, NER F-score, dependency parsing LAS, and text classification accuracy.
Citation
@misc{chang2020bilingual,
title={Towards Fully Bilingual Deep Language Modeling},
author={Chang et al. (2020)},
year={2020},
note={arXiv:2010.11639}
}
1---2name: finnish-nlp-eval3description: finnish-nlp-eval4---56# finnish-nlp-eval78> Towards Fully Bilingual Deep Language Modeling — Chang et al. (2020) (arXiv:2010.11639, 2020)910## What this evaluates1112Evaluates Finnish NLP capabilities across four core tasks: part-of-speech tagging, named entity recognition, dependency parsing, and text classification on domain-specific and out-of-domain corpora. It probes a model's ability to handle morphologically rich language and varying text registers.1314## Datasets1516- **Finnish NLP Benchmarks** — total ?; splits: train (-1), dev (-1), test (-1)1718## Metrics1920- `UPOS` — range: [0, 1]21 - Unambiguous Part-Of-Speech tag accuracy for POS tagging on Universal Dependencies treebanks.22- `F-score` — range: [0, 1]23 - Mention-level F-score for Named Entity Recognition using IOB annotations.24- `LAS` — range: [0, 1]25 - Labeled Attachment Score for dependency parsing (accuracy of predicted heads and relations).26- `Accuracy` — range: [0, 1]27 - Proportion of correctly classified documents for text classification.28- `Average` **(primary)** — range: [0, 1]29 - Algebraic mean of the four task scores (POS, NER, parsing, classification).3031## Input / output format3233**Input**: Tokenized Finnish text, typically truncated to at most 256 tokens.3435**Output**: Task-specific predictions: POS tags, IOB entity tags, dependency heads/relations, or document class labels.3637## Scoring recipe3839```python40pos_score = upos_accuracy(y_true, y_pred)41ner_score = mention_level_f1(y_true, y_pred)42parse_score = labeled_attachment_score(y_true, y_pred)43cls_score = document_accuracy(y_true, y_pred)44average = (pos_score + ner_score + parse_score + cls_score) / 445```4647## Common pitfalls4849- Uses gold segmentation for dependency parsing instead of predicted segmentation.50- Truncates all documents to 256 tokens, which may disadvantage models with longer context windows.51- Averages different metrics (UPOS, F1, LAS, Accuracy) without normalization, which can skew the final score.5253## Evidence (verbatim from paper)5455> The metric for POS tagging is UPOS, NER F-score, dependency parsing LAS, and text classification accuracy.5657## Citation5859```bibtex60@misc{chang2020bilingual,61 title={Towards Fully Bilingual Deep Language Modeling},62 author={Chang et al. (2020)},63 year={2020},64 note={arXiv:2010.11639}65}66```6768- arXiv: 2010.11639