Boolq Eval

This benchmark evaluates a model's ability to answer naturally occurring yes/no questions based on a provided passage. It probes complex inferential reasoning and non-factoid inference, requiring the model to go beyond simple keyword matching or shallow statistical features to determine entailment or contradiction between the question and the passage. Use when the user wants to benchmark on BoolQ, or asks about evaluating this task. Reports accuracy.

qhjqhj00 573c0e4 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/boolq-eval commit 573c0e47d2

Frequently asked questions

npx skillmds add qhjqhj00/boolq-eval