Mnli Anli Eval

Evaluates the out-of-domain generalization and robustness of NLI models trained under different data collection protocols. It measures how well models perform on held-out, genre-diverse, and adversarial benchmarks compared to in-domain validation performance. Use when the user wants to benchmark on MNLI-mismatched, ANLI, or asks about evaluating this task. Reports accuracy.

qhjqhj00 17cec2f 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mnli-anli-eval commit 17cec2f0e4

Frequently asked questions

npx skillmds add qhjqhj00/mnli-anli-eval