Towervision Multilingual Vl Eval

This evaluation probes the multilingual vision-language capabilities of models across text recognition, cultural understanding, multimodal translation, and video reasoning. It specifically tests cross-lingual generalization and cultural grounding in both image and video domains across high- and low-resource languages. Use when the user wants to benchmark on ALM-Bench, OCRBench, cc-OCR, TextVQA, CoMMuTE, Multi30K, ViMUL-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 86378ea 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/towervision-multilingual-vl-eval commit 86378ea1c2

Frequently asked questions

npx skillmds add qhjqhj00/towervision-multilingual-vl-eval