Audio Crowd Mt Eval

This protocol evaluates machine translation quality by comparing crowd-sourced human judgments of text-only outputs versus multimodal (text + audio) outputs. It probes whether audio-based assessments improve inter-rater consistency and reveal system-level differences through prosodic and expressive features unavailable in text. Use when the user wants to benchmark on WMT German-English, or asks about evaluating this task. Reports standardized score.

qhjqhj00 822d489 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/audio-crowd-mt-eval commit 822d4891e8

Frequently asked questions

npx skillmds add qhjqhj00/audio-crowd-mt-eval