Diffusion Instruction Tuning Eval

This protocol evaluates the vision-language alignment and zero-shot generalization capabilities of fine-tuned VLMs across diverse multimodal tasks. It measures how effectively aligning VLM cross-attention with diffusion model attention maps improves performance on document understanding, reasoning, real-world visual comprehension, and hallucination detection benchmarks. Use when the user wants to benchmark on AI2D, ChartQA, OCRBench, DocVQA, InfoVQA, MME, MMBench, ScienceQA, MMStar, MMMU, RealworldQA, SEED, HallucinationBench, POPE, VQAv2, OK-VQA, TextVQA, VizWiz, HatefulMemes, COCO, Flickr30K, WorldMedQA-V, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4da4bf9 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/diffusion-instruction-tuning-eval commit 4da4bf9cb8

Frequently asked questions

npx skillmds add qhjqhj00/diffusion-instruction-tuning-eval