Notes Bank Eval

Evaluates vision-language models on evidence-based visual question answering over unstructured, handwritten scientific notes. The benchmark probes a model's ability to localize relevant visual evidence via bounding boxes, classify content types, and generate natural language answers explicitly grounded in the visual input. Use when the user wants to benchmark on NoTeS-Bank, or asks about evaluating this task. Reports NDCG@5.

qhjqhj00 3795797 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/notes-bank-eval commit 379579718d

Frequently asked questions

npx skillmds add qhjqhj00/notes-bank-eval