Adversarial Nli Eval

Evaluates natural language inference models on adversarially crafted examples designed to expose reasoning brittleness and spurious pattern reliance. It probes whether models can generalize to novel, difficult inference cases that specifically target known model weaknesses across iterative rounds of human-and-model-in-the-loop data collection. Use when the user wants to benchmark on ANLI, or asks about evaluating this task. Reports accuracy.

qhjqhj00 44d1c5f 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/adversarial-nli-eval commit 44d1c5fc69

Frequently asked questions

npx skillmds add qhjqhj00/adversarial-nli-eval