Blimp Eval

This benchmark probes language models' sensitivity to grammatical acceptability contrasts across 12 linguistic phenomena. It evaluates whether models can reliably distinguish acceptable sentences from minimally ungrammatical ones, revealing strengths in morphological agreement and weaknesses in complex syntactic and semantic constraints. Use when the user wants to benchmark on BLiMP, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8abe6a2 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/blimp-eval commit 8abe6a2d6e

Frequently asked questions

npx skillmds add qhjqhj00/blimp-eval