Commonsense Retrieval Eval

This benchmark probes the commonsense reasoning capabilities of vision-language models by evaluating their ability to match images to text riddles (or vice versa) where the subject entity is replaced with a demonstrative pronoun. It specifically tests relational knowledge retrieval and generalization to unseen knowledge triples. Use when the user wants to benchmark on DANCE Diagnostic Set, or asks about evaluating this task. Reports Acc@50.

qhjqhj00 c8f7e31 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/commonsense-retrieval-eval commit c8f7e31e9a

Frequently asked questions

npx skillmds add qhjqhj00/commonsense-retrieval-eval