Kosmos1 Eval

Evaluates multimodal large language models on a comprehensive suite of perception-language, nonverbal reasoning, OCR-free text understanding, and web page comprehension tasks. It measures zero-shot and few-shot cross-modal transfer, in-context learning, and the ability to align visual perception with language generation without external tools or fine-tuning. Use when the user wants to benchmark on MS COCO Caption, Flickr30k, VQAv2, VizWiz, Raven IQ Test, Rendered SST-2, HatefulMemes, WebSRC, or asks about evaluating this task. Reports CIDEr, VQA accuracy.

qhjqhj00 08d332d 5.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kosmos1-eval commit 08d332dfd1

Frequently asked questions

npx skillmds add qhjqhj00/kosmos1-eval