topic-level-polarity-eval
SemEval-2015 Task 10: Sentiment Analysis in Twitter — Rosenthal et al. (2015) (SemEval-2015 / arXiv:1912.02387, 2015)
What this evaluates
Predicts the sentiment polarity associated with a specific topic within a tweet, requiring topic-aware sentiment classification and contextual disambiguation.
Datasets
- Twitter2015-test — total ?; splits: test (-1)
Metrics
macro-averaged F1(primary) — range: [0, 1]- Harmonic mean of precision and recall averaged across all classes (positive, negative, neutral).
Input / output format
Input: A tweet along with a target topic label.
Output: Sentiment polarity label for the specified topic.
Scoring recipe
prec = rec = 0
for class in ['pos', 'neg', 'neu']:
prec += precision(y_true, y_pred, class)
rec += recall(y_true, y_pred, class)
macro_prec = prec / 3
macro_rec = rec / 3
f1 = 2 * (macro_prec * macro_rec) / (macro_prec + macro_rec)
Common pitfalls
- Subtask C is significantly harder than B; top teams score 41-51 F1 vs 64-65 for B, despite similar class distributions.
- Performance drop is due to task complexity (topic-awareness), not class imbalance, as baseline F1 is similar (26.7% vs 30.3%).
Evidence (verbatim from paper)
The results for subtask C are shown in Table 12. ... achieved an F1 of 50.51, KLUEless with F1 = 45.48, and Whu-Nlp with F1 = 40.70
Citation
@misc{rosenthal2015semeval,
title={SemEval-2015 Task 10: Sentiment Analysis in Twitter},
author={Rosenthal et al. (2015)},
year={2015},
note={SemEval-2015 / arXiv:1912.02387}
}
- arXiv: 1912.02387