Mrag Bench Eval

Evaluates large vision-language models' ability to leverage retrieved visual knowledge versus textual knowledge across perspective and transformative change scenarios. Probes robustness to noisy retrieved images and measures how effectively models utilize visually augmented information compared to human baselines. Use when the user wants to benchmark on MRAG-Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6454818 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mrag-bench-eval commit 6454818d38

Frequently asked questions

npx skillmds add qhjqhj00/mrag-bench-eval