Dowis Eval

This benchmark evaluates instruction-following capabilities of speech-large language models (SLLMs) by comparing performance when prompted with text versus spoken audio across nine diverse tasks. It probes cross-lingual generalization, prompt style robustness, and the model's ability to handle both text and speech modalities for input and output. Use when the user wants to benchmark on DOWIS (Do What I Say), FLEURS, MCIF, YTSeg, or asks about evaluating this task. Reports WER.

qhjqhj00 f8ba28d 4.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dowis-eval commit f8ba28df5f

Frequently asked questions

npx skillmds add qhjqhj00/dowis-eval