Multimodal Instruction Eval

This protocol evaluates how concept- versus skill-targeted instruction selection strategies improve vision-language model performance under strict data budget constraints. It probes the model's zero-shot generalization across diverse tasks including VQA, OCR, spatial reasoning, and scientific understanding by aligning training data with the benchmark's dominant cognitive demand. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA (SQA-I), TextVQA, POPE, MME, MMBench (en), LLaVA-Bench, AI2D, OK-VQA, ST-VQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a8daa3b 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-instruction-eval commit a8daa3b0c5

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-instruction-eval