Sound Localization Eval

Evaluates a model's ability to localize sound sources in audio-visual pairs by predicting spatial response maps or bounding boxes. It measures how effectively the model aligns audio signals with visual regions containing the corresponding sound, particularly testing robustness to semantically similar but mismatched cross-modal pairs. Use when the user wants to benchmark on VGGSound, SoundNet-Flickr, VGG-SS, SoundNet-Flickr-Test, or asks about evaluating this task. Reports cIoU.

qhjqhj00 fb0c273 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sound-localization-eval commit fb0c273fab

Frequently asked questions

npx skillmds add qhjqhj00/sound-localization-eval