Textvr Retrieval Eval

Evaluates cross-modal video retrieval models that must jointly process visual context and scene text (OCR tokens) to match sentence queries with relevant videos. Probes the model's ability to read, comprehend, and align fine-grained text semantics with visual frames in real-world scenarios. Use when the user wants to benchmark on TextVR, or asks about evaluating this task. Reports R@K (Recall@K).

qhjqhj00 107d5e4 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/textvr-retrieval-eval commit 107d5e413f

Frequently asked questions

npx skillmds add qhjqhj00/textvr-retrieval-eval