Zero Shot Commonsense Eval

Evaluates zero-shot commonsense reasoning capabilities of language models using multiple-choice questions. It specifically probes how prompt engineering and probability calibration strategies affect accuracy across different model sizes and architectures. Use when the user wants to benchmark on CommonsenseQA, COPA, OpenBookQA, PIQA, Social IQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7a8cef1 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/zero-shot-commonsense-eval commit 7a8cef186d

Frequently asked questions

npx skillmds add qhjqhj00/zero-shot-commonsense-eval