Save Video Text Retrieval Eval

Evaluates a model's ability to retrieve relevant videos given a natural language query in an audio-visual setting. It specifically probes how well speech-aware representations and early vision-audio alignment improve cross-modal matching accuracy across diverse video-text benchmarks. Use when the user wants to benchmark on MSRVTT-9k, MSRVTT-7k, VATEX, Charades, LSMDC, or asks about evaluating this task. Reports SumR.

qhjqhj00 50982c1 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/save-video-text-retrieval-eval commit 50982c1133

Frequently asked questions

npx skillmds add qhjqhj00/save-video-text-retrieval-eval