Literaryqa Eval

Evaluates long-context language models on their ability to answer complex, abstractive questions about entire literary works. It probes narrative event understanding and semantic correctness, measuring how well generated answers align with human judgment rather than just matching reference strings. Use when the user wants to benchmark on LiteraryQA, or asks about evaluating this task. Reports ROUGE-L.

qhjqhj00 7ff2903 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/literaryqa-eval commit 7ff2903e81

Frequently asked questions

npx skillmds add qhjqhj00/literaryqa-eval