dynamic-superb-eval
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech — Huang et al. (2023) (arXiv:2309.09510, 2023)
What this evaluates
Evaluates instruction-tuned speech models on their ability to perform diverse speech and audio tasks using natural language instructions. It probes zero-shot generalization by testing performance on seen versus unseen tasks and instructions across six dimensions: content, speaker, semantics, degradation, paralinguistics, and audio.
Datasets
- Dynamic-SUPERB — total 55; splits: test (55); repo https://github.com/dynamic-superb/dynamic-superb
Metrics
accuracy(primary) — range: percent- Standard classification accuracy: the percentage of instances where the model's generated output exactly matches the ground-truth label among the provided options.
Input / output format
Input: Audio waveform paired with a natural language instruction that specifies the task and provides a set of candidate answer options.
Output: A generated text string corresponding to one of the candidate options provided in the instruction.
Scoring recipe
correct = 0
for pred, gold in zip(predictions, gold_labels):
if pred.strip().lower() == gold.strip().lower():
correct += 1
accuracy = (correct / len(predictions)) * 100
Common pitfalls
- Models may generate outputs not present in the candidate options, causing accuracy to drop below the random baseline.
- Models often rely on superficial instruction patterns (bag-of-words) rather than semantic understanding, leading to poor generalization on unseen tasks.
- The random baseline is constructed by sampling from the label distribution of each task, not a uniform random guess.
Evidence (verbatim from paper)
Table 2 presents the evaluation results for each baseline on seen tasks in Dynamic-SUPERB. Due to space constraints, we report the average accuracy for each dimension.
Citation
@misc{huang2023dynamicsuperb,
title={Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech},
author={Huang et al. (2023)},
year={2023},
note={arXiv:2309.09510}
}
- arXiv: 2309.09510