Mtr Duplexbench Eval

This benchmark evaluates Full-Duplex Speech Language Models (FD-SLMs) on their ability to sustain performance across multi-round conversations. It probes dialogue quality, conversational dynamics (turn-taking, interruptions, pauses, background speech), instruction following, and safety, specifically measuring how these capabilities degrade or hold up as interaction rounds increase. Use when the user wants to benchmark on MTR-DuplexBench, Llama Question, AdvBench, or asks about evaluating this task. Reports Success Rate (%).

qhjqhj00 dd5f81d 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mtr-duplexbench-eval commit dd5f81da84

Frequently asked questions

npx skillmds add qhjqhj00/mtr-duplexbench-eval