Tempo Sum Eval

Evaluates text summarization models' temporal generalization by testing on datasets split by publication date, specifically probing how well models handle knowledge-conflicting future articles versus in-distribution past data. Use when the user wants to benchmark on BBC, CNN, or asks about evaluating this task. Reports FactCC.

qhjqhj00 5b7cb08 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tempo-sum-eval commit 5b7cb085db

Frequently asked questions

npx skillmds add qhjqhj00/tempo-sum-eval