Caselawqa Eval

This benchmark probes a model's ability to perform fine-grained legal text classification and annotation. It tests whether models can accurately extract specific legal features, such as precedent alteration, issue areas, or ideological valence, from lengthy court opinions using multiple-choice prompts. Use when the user wants to benchmark on CaselawQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7a1e0bc 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/caselawqa-eval commit 7a1e0bc662

Frequently asked questions

npx skillmds add qhjqhj00/caselawqa-eval