# Smtce Eval

> smtce-eval

- Skill: `qhjqhj00/smtce-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/smtce-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/smtce-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/smtce-eval

---


# smtce-eval

> SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese — Luan Thanh Nguyen et al. (2022) (arXiv:2209.10482, 2022)

## What this evaluates

Evaluates Vietnamese social media text classification across four tasks: constructive speech detection, complaint detection, emotion recognition, and hate speech detection. It probes the ability of monolingual versus multilingual BERT-based models to handle low-resource, domain-specific Vietnamese text with varying preprocessing requirements.

## Datasets

- **VSMEC** — total ?; splits: test (-1)
- **ViCTSD** — total ?; splits: test (-1)
- **ViOCD** — total ?; splits: test (-1)
- **ViHSD** — total ?; splits: test (-1)

## Metrics

- `macro-average F1 score` **(primary)** — range: percent
  - Compute Precision and Recall per class, then F1 per class as 2*(P*R)/(P+R). Average these per-class F1 scores across all classes.

## Input / output format

**Input**: Vietnamese social media text (comments/posts), preprocessed according to dataset-specific rules (e.g., emoji preservation for emotion tasks, number removal for others, tokenization via VnCoreNLP or FAIRSeq).

**Output**: Predicted class label for each text instance.

## Scoring recipe

```python
def compute_macro_f1(y_true, y_pred, num_classes):
    f1_scores = []
    for c in range(num_classes):
        tp = sum(1 for yt, yp in zip(y_true, y_pred) if yt == c and yp == c)
        fp = sum(1 for yt, yp in zip(y_true, y_pred) if yt != c and yp == c)
        fn = sum(1 for yt, yp in zip(y_true, y_pred) if yt == c and yp != c)
        precision = tp / (tp + fp) if (tp + fp) > 0 else 0
        recall = tp / (tp + fn) if (tp + fn) > 0 else 0
        f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0
        f1_scores.append(f1)
    return sum(f1_scores) / num_classes
```

## Common pitfalls

- Datasets are highly imbalanced, so accuracy is misleading; macro-F1 is required to fairly weight minority classes.
- Preprocessing must be task-specific (e.g., emojis must be preserved for emotion classification but removed for other tasks); applying uniform preprocessing degrades performance.
- Multilingual models often underperform monolingual ones due to domain mismatch and limited Vietnamese social media exposure in pretraining.

## Evidence (verbatim from paper)

> In the text classification task, we have different metrics suitable for specific datasets and problems. Because most datasets in this study are imbalanced and according to the choice of dataset authors, we choose the macro-average F1 score to evaluate the performances of models on the datasets. To compute the macro-average F1 score, we first calculate the F1 score per class in the dataset by the formula (1). F1 score = 2 * (Precision * Recall) / (Precision + Recall). After achieving the F1 scores of all classes, we compute the macro-average F1 score by calculating the average F1 score as shown in formula (2). Macro F1 score = sum(F1 scores) / Number of classes.

## Citation

```bibtex
@misc{nguyen2022smtce,
  title={SMTCE: A Social Media Text Classification Evaluation Benchmark and BERTology Models for Vietnamese},
  author={Luan Thanh Nguyen et al. (2022)},
  year={2022},
  note={arXiv:2209.10482}
}
```

- arXiv: 2209.10482

