Wep Verbalization Validity Eval

Evaluates neural language models' ability to understand Words of Estimative Probability (WEP) by testing their capacity to distinguish valid from invalid probabilistic verbalizations and perform logical consistency checks in probabilistic reasoning. Use when the user wants to benchmark on WEP Reasoning 1 hop, WEP Reasoning 2 hops, WEP-UNLI, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7f51d71 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wep-verbalization-validity-eval commit 7f51d713b2

Frequently asked questions

npx skillmds add qhjqhj00/wep-verbalization-validity-eval