propaganda-detection-eval
SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles — Da San Martino et al. (2020) (arXiv:2009.02696, 2020)
What this evaluates
Detects propaganda techniques in news articles through a two-stage pipeline: identifying text spans containing propaganda (Span Identification) and classifying the specific rhetorical technique used within those spans (Technique Classification).
Datasets
- SemEval-2020 Task 11 Propaganda Detection — total ?; splits: test (-1), development (-1)
Metrics
F1(primary) — range: [0, 1]- Harmonic mean of precision and recall. For Span Identification, computed at the character level over identified spans. For Technique Classification, macro-averaged across 14 propaganda technique classes.
Input / output format
Input: News article text with character-level annotations for propaganda spans and their corresponding technique labels.
Output: For Span Identification: a list of character spans flagged as propaganda. For Technique Classification: a classification label from 14 predefined propaganda techniques for each identified span.
Scoring recipe
def compute_f1(pred_spans, gold_spans, pred_labels, gold_labels):
# Span Identification (character-level)
tp = len(set(pred_spans) & set(gold_spans))
fp = len(set(pred_spans) - set(gold_spans))
fn = len(set(gold_spans) - set(pred_spans))
prec = tp / (tp + fp) if (tp + fp) > 0 else 0
rec = tp / (tp + fn) if (tp + fn) > 0 else 0
f1_si = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
# Technique Classification (macro-averaged)
f1_tc = macro_f1(pred_labels, gold_labels)
return f1_si, f1_tc
Common pitfalls
- Overfitting to the development set can cause significant performance drops on the test set, as observed for teams syrapropa and PALI.
- Combining systems via union increases recall but decreases precision, while intersection does the opposite; majority voting offers a balance.
- Rare propaganda techniques (e.g., Straw man, Bandwagon) are hard to classify, skewing overall macro-F1 performance.
Evidence (verbatim from paper)
Figure 6: F1 performance for the technique classification subtask when combining up to top-20 systems with majority voting. The plots show the overall performance (top left) as well as for each of the 14 classes.
Citation
@misc{dasanmartino2020semeval,
title={SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles},
author={Da San Martino et al. (2020)},
year={2020},
note={arXiv:2009.02696}
}
- arXiv: 2009.02696