Global Piqa Eval

This benchmark probes physical commonsense reasoning by testing whether models can distinguish correct from incorrect solutions to everyday physical tasks. It specifically evaluates cultural and linguistic grounding by using items constructed natively in 116 language varieties, avoiding translation artifacts that often skew multilingual evaluations. Use when the user wants to benchmark on Global PIQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 fb2be81 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/global-piqa-eval commit fb2be81211

Frequently asked questions

npx skillmds add qhjqhj00/global-piqa-eval