Flava Multimodal Vision Nlp Eval

Evaluates a unified vision-language foundation model across 35 downstream tasks spanning vision classification, natural language understanding, and multimodal reasoning/retrieval. It probes the model's ability to generalize from joint unimodal and multimodal pretraining to zero-shot and fine-tuned downstream settings. Use when the user wants to benchmark on GLUE (MNLI, CoLA, MRPC, QQP, SST-2, QNLI, RTE, STS-B), 22 Vision Datasets (ImageNet, Food101, CIFAR10, CIFAR100, Cars, Aircraft, DTD, Pets, Caltech101, Flowers102, MNIST, STL10, EuroSAT, GTSRB, KITTI, PCAM, UCF101, CLEVR, FER 2013, SUN397, SST, Country211), VQAv2, SNLI-VE, Hateful Memes, Flickr30K, COCO, or asks about evaluating this task. Reports accuracy.

qhjqhj00 57f91af 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/flava-multimodal-vision-nlp-eval commit 57f91af9b9

Frequently asked questions

npx skillmds add qhjqhj00/flava-multimodal-vision-nlp-eval