Sagi Eval

This benchmark evaluates speech large language models across five hierarchical levels of understanding, ranging from basic automatic speech recognition and language identification to paralinguistic perception (pitch, volume, emotion), abstract acoustic reasoning (medical cough analysis), and creative/agentic tasks (spoken English coaching). It probes the model's ability to process raw audio, follow instructions, and extract both semantic and non-semantic acoustic features. Use when the user wants to benchmark on SAGI Benchmark, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 71f6a7a 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sagi-eval commit 71f6a7a1d9

Frequently asked questions

npx skillmds add qhjqhj00/sagi-eval