harmonysset-eval
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization — Zhou et al. (2025) (arXiv:2503.01725, 2025)
What this evaluates
Evaluates multimodal large language models' ability to align video and music across four dimensions: rhythmic synchronization, thematic coherence, emotional congruence, and cultural relevance. It tests both open-ended descriptive reasoning and multiple-choice selection to measure temporal and semantic alignment capabilities.
Datasets
- HarmonySet — total 48328; splits: test (-1)
Metrics
HarmonySet-OE Score(primary) — range: other- Average score across four dimensions: Rhythm & Synchronization (R&S), Thematic (T), Emotional (E), and Cultural (C). Scores are reported on a continuous scale (values range ~3.0–5.6 in tables).
HarmonySet-MC Accuracy— range: percent- Percentage of correctly selected multiple-choice answers across the four dimensions (R&S, T, E, C).
Input / output format
Input: Video clip (sampled at 16 frames for open-source models, 1 fps for Gemini 1.5 Pro) paired with audio/music, accompanied by a prompt/question for either open-ended generation or multiple-choice selection.
Output: Open-ended textual analysis for HarmonySet-OE, or a single selected option for HarmonySet-MC.
Scoring recipe
# For HarmonySet-MC
accuracy = sum(1 for pred, gold in zip(predictions, gold_labels) if pred == gold) / len(gold_labels) * 100
# For HarmonySet-OE
# Scores are assigned per dimension (R&S, T, E, C) by human annotators or automated scorers.
# Final OE Score = mean(score_R&S, score_T, score_E, score_C)
Common pitfalls
- Frame sampling rate critically impacts performance; using 64 frames on short videos (<1 min) degrades scores due to redundancy/overfitting, while 16-32 frames are optimal.
- Zero-shot evaluation requires identical prompts across all models; inconsistent prompting or input frame counts invalidate cross-model comparisons.
- Cultural relevance and temporal synchronization are particularly difficult dimensions where even state-of-the-art models significantly lag behind human performance.
Evidence (verbatim from paper)
We conduct the evaluation on Gemini 1.5 Pro and state-of-the-art open-source video-audio MLLMs, including VideoLLaMA2 and video-SALMONN. For a fair comparison, we adopt the zero-shot setting to infer HarmonySet-OE questions with all MLLMs based on the same prompt. In the experiments presented in Table [2], we used a consistent 16 frames for the video input of open-source models for both inference and fine-tuning. Table [5] shows the main results on HarmonySet-OE. Table [5]: Human and model performance on HarmonySet-MC. While VideoLLaMA2 tuned on HarmonySet surpasses Gemini-1.5 Pro in certain aspects, it still falls short of human performance, highlighting both the challenging nature of our task and the limitations of current models.
Citation
@misc{zhou2025harmonysset,
title={HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization},
author={Zhou et al. (2025)},
year={2025},
note={arXiv:2503.01725}
}
- arXiv: 2503.01725