Coliee Task4 Eval

Evaluates the ability of large language models to perform legal textual entailment, specifically measuring how model accuracy changes over time based on the year of the Japanese statute law data used. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f3b6e03 2.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/coliee-task4-eval commit f3b6e03b3f

Frequently asked questions

npx skillmds add qhjqhj00/coliee-task4-eval