Audiocrag Eval

Evaluates speech-in-speech-out dialogue systems on their ability to accurately answer spoken queries using external tools, measuring both answer correctness and system latency under streaming versus open-book settings. Use when the user wants to benchmark on AudioCRAG, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f80ca84 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/audiocrag-eval commit f80ca84221

Frequently asked questions

npx skillmds add qhjqhj00/audiocrag-eval