message-level-polarity-eval
SemEval-2015 Task 10: Sentiment Analysis in Twitter — Rosenthal et al. (2015) (SemEval-2015 / arXiv:1912.02387, 2015)
What this evaluates
Determines the overall sentiment polarity of an entire tweet, addressing class imbalance, slang, and informal social media text.
Datasets
- Twitter2015-test — total ?; splits: test (-1)
Metrics
macro-averaged F1(primary) — range: [0, 1]- Harmonic mean of precision and recall averaged across all classes (positive, negative, neutral).
Input / output format
Input: A full tweet.
Output: Sentiment polarity label (positive, negative, or neutral).
Scoring recipe
prec = rec = 0
for class in ['pos', 'neg', 'neu']:
prec += precision(y_true, y_pred, class)
rec += recall(y_true, y_pred, class)
macro_prec = prec / 3
macro_rec = rec / 3
f1 = 2 * (macro_prec * macro_rec) / (macro_prec + macro_rec)
Common pitfalls
- Baseline majority class F1 is ~30.3%, so models must significantly outperform simple majority voting.
- Sarcastic tweets are a subset of the test set and evaluated separately; models trained on non-sarcastic data often degrade on this subset.
Evidence (verbatim from paper)
The results for subtask B are shown in Table 11. ... with an F1 of 64.84, unitn with 64.59, lislif with 64.27, and INESC-ID with 64.17.
Citation
@misc{rosenthal2015semeval,
title={SemEval-2015 Task 10: Sentiment Analysis in Twitter},
author={Rosenthal et al. (2015)},
year={2015},
note={SemEval-2015 / arXiv:1912.02387}
}
- arXiv: 1912.02387