Scienceqa Eval

Evaluates multimodal reasoning and scientific question answering by requiring models to process questions, images, and context to select correct multiple-choice answers. It also probes chain-of-thought reasoning capabilities by measuring the quality of generated explanations and lectures. Use when the user wants to benchmark on ScienceQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b38646b 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scienceqa-eval commit b38646b439

Frequently asked questions

npx skillmds add qhjqhj00/scienceqa-eval