# Mtat Nsynth Fma Eval

> mtat-nsynth-fma-eval

- Skill: `qhjqhj00/mtat-nsynth-fma-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/mtat-nsynth-fma-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/mtat-nsynth-fma-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/mtat-nsynth-fma-eval

---


# mtat-nsynth-fma-eval

> Multi-Source Contrastive Learning from Musical Audio — Garoufis et al. (2023) (arXiv:2302.07077, 2023)

## What this evaluates

Evaluates the quality of self-supervised audio representations on three downstream tasks: music auto-tagging, instrument family classification, and music genre classification. It probes how well contrastive learning objectives capture semantic and structural musical features.

## Datasets

- **MTAT** — total ?; splits: train (-1), val (-1), test (-1); repo https://github.com/jongpillee/music_dataset_split
- **NSynth** — total 305979; splits: train (289205), val (12678), test (4096)
- **FMA** — total ?; splits: train (-1), val (-1), test (-1)

## Metrics

- `ROC-AUC` **(primary)** — range: [0, 1]
  - Macro average Receiver Operating Characteristic Area Under the Curve computed over all tags for multi-label auto-tagging.
- `PR-AUC` — range: [0, 1]
  - Macro average Precision-Recall Area Under the Curve computed over all tags for multi-label auto-tagging.
- `WA (%)` — range: percent
  - Weighted Accuracy percentage computed over classes for instrument family and genre classification.

## Input / output format

**Input**: 1-second audio segments extracted from tracks.

**Output**: Per-segment prediction probabilities aggregated via averaging to produce instance-level predictions for multi-label tags or class labels.

## Scoring recipe

```python
def evaluate(segments, y_true, is_multilabel):
    # Aggregate per-segment predictions to instance level
    preds = average_segment_predictions(segments)
    
    if is_multilabel:
        roc_auc = macro_roc_auc(y_true, preds)
        pr_auc = macro_pr_auc(y_true, preds)
        return roc_auc, pr_auc
    else:
        wa = weighted_accuracy(y_true, preds)
        return wa

# Report mean over 5 independent runs with different random seeds
```

## Common pitfalls

- Models receive 1-sec segments, but evaluation aggregates predictions per instance (track) by averaging probabilities, not by majority vote or argmax.
- Metrics are macro-averaged over all tags/classes, not micro-averaged or per-class.
- Results are averaged over 5 independent runs with different random seeds, not a single split.

## Evidence (verbatim from paper)

> Since music auto-tagging constitutes a multi-instance multi-label task, we use the categorical cross-entropy as its loss function, and report on the macro average ROC-AUC and PR-AUC values over all tags. On the other hand, for both instrument family classification and music genre classification, the binary cross-entropy is used as the loss function and the weighted accuracy (WA, %) as the evaluation metric. During evaluation, all models receive 1-sec audio segments, and aggregate the predictions for each instance via averaging the per-segment predictions. We further note that all downstream classification experiments were repeated 5 times, and we report on their average scores.

## Citation

```bibtex
@misc{garoufis2023multisourcecontrastive,
  title={Multi-Source Contrastive Learning from Musical Audio},
  author={Garoufis et al. (2023)},
  year={2023},
  note={arXiv:2302.07077}
}
```

- arXiv: 2302.07077

