# Harmonysset Eval

> harmonysset-eval

- Skill: `qhjqhj00/harmonysset-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/harmonysset-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/harmonysset-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/harmonysset-eval

---


# harmonysset-eval

> HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization — Zhou et al. (2025) (arXiv:2503.01725, 2025)

## What this evaluates

Evaluates multimodal large language models' ability to align video and music across four dimensions: rhythmic synchronization, thematic coherence, emotional congruence, and cultural relevance. It tests both open-ended descriptive reasoning and multiple-choice selection to measure temporal and semantic alignment capabilities.

## Datasets

- **HarmonySet** — total 48328; splits: test (-1)

## Metrics

- `HarmonySet-OE Score` **(primary)** — range: other
  - Average score across four dimensions: Rhythm & Synchronization (R&S), Thematic (T), Emotional (E), and Cultural (C). Scores are reported on a continuous scale (values range ~3.0–5.6 in tables).
- `HarmonySet-MC Accuracy` — range: percent
  - Percentage of correctly selected multiple-choice answers across the four dimensions (R&S, T, E, C).

## Input / output format

**Input**: Video clip (sampled at 16 frames for open-source models, 1 fps for Gemini 1.5 Pro) paired with audio/music, accompanied by a prompt/question for either open-ended generation or multiple-choice selection.

**Output**: Open-ended textual analysis for HarmonySet-OE, or a single selected option for HarmonySet-MC.

## Scoring recipe

```python
# For HarmonySet-MC
accuracy = sum(1 for pred, gold in zip(predictions, gold_labels) if pred == gold) / len(gold_labels) * 100

# For HarmonySet-OE
# Scores are assigned per dimension (R&S, T, E, C) by human annotators or automated scorers.
# Final OE Score = mean(score_R&S, score_T, score_E, score_C)
```

## Common pitfalls

- Frame sampling rate critically impacts performance; using 64 frames on short videos (<1 min) degrades scores due to redundancy/overfitting, while 16-32 frames are optimal.
- Zero-shot evaluation requires identical prompts across all models; inconsistent prompting or input frame counts invalidate cross-model comparisons.
- Cultural relevance and temporal synchronization are particularly difficult dimensions where even state-of-the-art models significantly lag behind human performance.

## Evidence (verbatim from paper)

> We conduct the evaluation on Gemini 1.5 Pro and state-of-the-art open-source video-audio MLLMs, including VideoLLaMA2 and video-SALMONN. For a fair comparison, we adopt the zero-shot setting to infer HarmonySet-OE questions with all MLLMs based on the same prompt. In the experiments presented in Table [2], we used a consistent 16 frames for the video input of open-source models for both inference and fine-tuning. Table [5] shows the main results on HarmonySet-OE. Table [5]: Human and model performance on HarmonySet-MC. While VideoLLaMA2 tuned on HarmonySet surpasses Gemini-1.5 Pro in certain aspects, it still falls short of human performance, highlighting both the challenging nature of our task and the limitations of current models.

## Citation

```bibtex
@misc{zhou2025harmonysset,
  title={HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization},
  author={Zhou et al. (2025)},
  year={2025},
  note={arXiv:2503.01725}
}
```

- arXiv: 2503.01725

