Vlm Deflection Bench Eval

This benchmark evaluates the ability of large vision-language models to correctly answer knowledge-based visual questions while properly deferring when evidence is missing or hallucinating when faced with noisy or conflicting retrieval contexts. It disentangles parametric memorization from retrieval robustness across four controlled scenarios. Use when the user wants to benchmark on VLM-DeflectionBench, or asks about evaluating this task. Reports Deflection Rate.

qhjqhj00 9e1e255 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vlm-deflection-bench-eval commit 9e1e255154

Frequently asked questions

npx skillmds add qhjqhj00/vlm-deflection-bench-eval