Imagenethink250k Eval

Evaluates vision-language models' ability to generate structured, step-by-step reasoning (thinking tokens) and final answers for multimodal inputs. It probes reasoning coherence, logical progression, and alignment with reference synthetic reasoning traces. Use when the user wants to benchmark on ImageNet-Think-250K, or asks about evaluating this task. Reports BERTScore.

qhjqhj00 8c20a83 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/imagenethink250k-eval commit 8c20a83f88

Frequently asked questions

npx skillmds add qhjqhj00/imagenethink250k-eval