Vggsounder Eval

Evaluates audio-visual foundation and embedding models on multi-label video classification, probing their ability to recognize sound and visual events across different input modalities. It specifically measures modality alignment, unimodal versus multimodal performance, and susceptibility to distraction from irrelevant background audio or static visuals. Use when the user wants to benchmark on VGGSounder, or asks about evaluating this task. Reports F1-score.

qhjqhj00 8609607 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vggsounder-eval commit 86096078ef

Frequently asked questions

npx skillmds add qhjqhj00/vggsounder-eval