Speechjudge Eval

This benchmark evaluates how well automated metrics and large audio-language models can judge speech naturalness and detect deepfakes compared to human preferences. It probes the alignment of computational quality scores with human perceptual judgments across multiple languages and TTS models. Use when the user wants to benchmark on SpeechJudge-Eval, or asks about evaluating this task. Reports AudioLLM Pairwise Accuracy.

qhjqhj00 33fc0cd 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/speechjudge-eval commit 33fc0cdafa

Frequently asked questions

npx skillmds add qhjqhj00/speechjudge-eval