Av Speech Enhancement Eval

This benchmark evaluates a model's ability to isolate a target speaker's voice from multi-talker audio environments using only lip-region video inputs. It probes audio-visual speech enhancement, testing how well the network predicts magnitude and phase masks to suppress interference and noise while preserving speech intelligibility and perceptual quality. Use when the user wants to benchmark on LRS2, VoxCeleb2, or asks about evaluating this task. Reports PESQ.

qhjqhj00 41fae6d 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/av-speech-enhancement-eval commit 41fae6d438

Frequently asked questions

npx skillmds add qhjqhj00/av-speech-enhancement-eval