mmla-eval
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark — Zhang et al. (2025) (arXiv:2504.16427, 2025)
What this evaluates
Evaluates the ability of LLMs and MLLMs to perform multimodal language analysis across six high-level semantic dimensions: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. It probes cross-modal reasoning and cognitive-level semantic understanding in conversational contexts using aligned text and video utterances.
Datasets
- MIntRec — total ?; splits: train (-1), val (-1), test (-1)
- MIntRec2.0 — total ?; splits: train (-1), val (-1), test (-1)
- MELD — total ?; splits: train (-1), val (-1), test (-1)
- IEMOCAP — total ?; splits: train (-1), val (-1), test (-1)
- MOSI — total ?; splits: train (-1), val (-1), test (-1)
- CH-SIMS v2.0 — total ?; splits: train (-1), val (-1), test (-1)
- UR-FUNNY-v2 — total ?; splits: train (-1), val (-1), test (-1)
- MUStARD — total ?; splits: train (-1), val (-1), test (-1)
Metrics
accuracy (ACC) (primary) — range: [0, 1]
- The proportion of correctly predicted labels out of the total number of instances. Calculated as (True Positives + True Negatives) / Total Instances.
weighted F1-score (WF1) — range: [0, 1]
- The harmonic mean of precision and recall, weighted by the number of true instances for each class.
weighted precision (WP) — range: [0, 1]
- The ratio of true positive predictions to the total number of positive predictions, weighted by class support.
macro F1-score (F1) — range: [0, 1]
- The unweighted mean of F1-scores calculated for each class independently.
recall (R) — range: [0, 1]
- The ratio of true positive predictions to the total number of actual positives.
precision (P) — range: [0, 1]
- The ratio of true positive predictions to the total number of predicted positives.
Input / output format
Input: Aligned text and video data at the utterance level for speakers.
Output: Classification label corresponding to the target semantic dimension (e.g., intent, emotion, dialogue act, sentiment, speaking style, or communication behavior).
Scoring recipe
def compute_accuracy(predictions, gold_labels):
correct = sum(1 for p, g in zip(predictions, gold_labels) if p == g)
return correct / len(gold_labels)
def compute_f1(precision, recall):
if precision + recall == 0:
return 0.0
return 2 * (precision * recall) / (precision + recall)
Common pitfalls
- Models must process both text and video modalities; evaluating only on text ignores the multimodal alignment requirement.
- The benchmark covers six distinct semantic dimensions, so results should be reported per dimension rather than aggregated blindly.
- Zero-shot performance is notably low; fine-tuning (SFT/IT) is required to reach 60-70% accuracy, so comparing only zero-shot baselines is misleading.
Evidence (verbatim from paper)
We employ six commonly used metrics: accuracy (ACC), weighted F1-score (WF1), weighted precision (WP), macro F1-score (F1), recall (R), and precision (P) for evaluation, as suggested in the literature. In particular, we report the primary results of ACC in this paper, with additional results for the remaining metrics provided in the Appendices.
Citation
@misc{zhang2025mmla,
title={Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark},
author={Zhang et al. (2025)},
year={2025},
note={arXiv:2504.16427}
}
1---2name: mmla-eval3description: mmla-eval4---56# mmla-eval78> Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark — Zhang et al. (2025) (arXiv:2504.16427, 2025)910## What this evaluates1112Evaluates the ability of LLMs and MLLMs to perform multimodal language analysis across six high-level semantic dimensions: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. It probes cross-modal reasoning and cognitive-level semantic understanding in conversational contexts using aligned text and video utterances.1314## Datasets1516- **MIntRec** — total ?; splits: train (-1), val (-1), test (-1)17- **MIntRec2.0** — total ?; splits: train (-1), val (-1), test (-1)18- **MELD** — total ?; splits: train (-1), val (-1), test (-1)19- **IEMOCAP** — total ?; splits: train (-1), val (-1), test (-1)20- **MOSI** — total ?; splits: train (-1), val (-1), test (-1)21- **CH-SIMS v2.0** — total ?; splits: train (-1), val (-1), test (-1)22- **UR-FUNNY-v2** — total ?; splits: train (-1), val (-1), test (-1)23- **MUStARD** — total ?; splits: train (-1), val (-1), test (-1)2425## Metrics2627- `accuracy (ACC)` **(primary)** — range: [0, 1]28 - The proportion of correctly predicted labels out of the total number of instances. Calculated as (True Positives + True Negatives) / Total Instances.29- `weighted F1-score (WF1)` — range: [0, 1]30 - The harmonic mean of precision and recall, weighted by the number of true instances for each class.31- `weighted precision (WP)` — range: [0, 1]32 - The ratio of true positive predictions to the total number of positive predictions, weighted by class support.33- `macro F1-score (F1)` — range: [0, 1]34 - The unweighted mean of F1-scores calculated for each class independently.35- `recall (R)` — range: [0, 1]36 - The ratio of true positive predictions to the total number of actual positives.37- `precision (P)` — range: [0, 1]38 - The ratio of true positive predictions to the total number of predicted positives.3940## Input / output format4142**Input**: Aligned text and video data at the utterance level for speakers.4344**Output**: Classification label corresponding to the target semantic dimension (e.g., intent, emotion, dialogue act, sentiment, speaking style, or communication behavior).4546## Scoring recipe4748```python49def compute_accuracy(predictions, gold_labels):50 correct = sum(1 for p, g in zip(predictions, gold_labels) if p == g)51 return correct / len(gold_labels)5253def compute_f1(precision, recall):54 if precision + recall == 0:55 return 0.056 return 2 * (precision * recall) / (precision + recall)57```5859## Common pitfalls6061- Models must process both text and video modalities; evaluating only on text ignores the multimodal alignment requirement.62- The benchmark covers six distinct semantic dimensions, so results should be reported per dimension rather than aggregated blindly.63- Zero-shot performance is notably low; fine-tuning (SFT/IT) is required to reach 60-70% accuracy, so comparing only zero-shot baselines is misleading.6465## Evidence (verbatim from paper)6667> We employ six commonly used metrics: accuracy (ACC), weighted F1-score (WF1), weighted precision (WP), macro F1-score (F1), recall (R), and precision (P) for evaluation, as suggested in the literature. In particular, we report the primary results of ACC in this paper, with additional results for the remaining metrics provided in the Appendices.6869## Citation7071```bibtex72@misc{zhang2025mmla,73 title={Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark},74 author={Zhang et al. (2025)},75 year={2025},76 note={arXiv:2504.16427}77}78```7980- arXiv: 2504.16427