Commonsenseqa Eval

This benchmark evaluates a model's ability to answer multiple-choice questions that require real-world commonsense knowledge. It specifically probes whether models can distinguish a correct answer from semantically plausible but factually incorrect distractors based on spatial, causal, or physical reasoning. Use when the user wants to benchmark on CommonsenseQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 cf8ccfd 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/commonsenseqa-eval commit cf8ccfd648

Frequently asked questions

npx skillmds add qhjqhj00/commonsenseqa-eval