Ar Bench Eval

Evaluates large language models on appellate review tasks for criminal judgments, specifically detecting, classifying, and correcting legal errors in finalized court decisions. It probes fine-grained legal reasoning, diagnostic accuracy, and the ability to generate legally valid corrections. Use when the user wants to benchmark on AR-Bench, or asks about evaluating this task. Reports Accuracy (Acc), Macro F1 (MaF1).

qhjqhj00 0ea12e8 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ar-bench-eval commit 0ea12e8db7

Frequently asked questions

npx skillmds add qhjqhj00/ar-bench-eval