Eraser Benchmark Eval

Evaluates NLP models' ability to generate faithful, task-appropriate rationales for predictions, measuring both alignment with human annotations and causal faithfulness via token perturbation. Use when the user wants to benchmark on Movies, FEVER, CoS-E, eSNLI, or asks about evaluating this task. Reports AUPRC.

qhjqhj00 5cc7184 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/eraser-benchmark-eval commit 5cc7184eaf

Frequently asked questions

npx skillmds add qhjqhj00/eraser-benchmark-eval