Wikicatsum Rouge Eval

Evaluates abstractive multi-document summarization models on their ability to generate coherent, content-adequate summaries across three domains (Company, Film, Animal). The protocol measures lexical and sentence-level overlap between generated summaries and reference summaries using ROUGE metrics, while also contextualizing scores against a baseline overlap between input documents and summaries. Use when the user wants to benchmark on WIKICATSUM, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L.

qhjqhj00 7c8010c 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wikicatsum-rouge-eval commit 7c8010c3f1

Frequently asked questions

npx skillmds add qhjqhj00/wikicatsum-rouge-eval