Feedsum Eval

This benchmark evaluates how well LLM-generated feedback aligns with human preferences for text summarization, and tests whether preference learning (DPO) using multi-dimensional feedback improves summary quality over supervised fine-tuning. Use when the user wants to benchmark on FeedSum, or asks about evaluating this task. Reports Spearman correlation.

qhjqhj00 ab39538 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/feedsum-eval commit ab39538d56

Frequently asked questions

npx skillmds add qhjqhj00/feedsum-eval