Longreward Eval

This evaluation protocol assesses the long-context understanding, instruction-following, and faithfulness capabilities of LLMs. It combines automated AI-judged scoring on long and short-context benchmarks with human preference alignment tests to validate the effectiveness of the LongReward training method. Use when the user wants to benchmark on LongBench, LongBench-Chat, MT-Bench, AlpacaEval2, or asks about evaluating this task. Reports GPT-4o rating.

qhjqhj00 4b986c8 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/longreward-eval commit 4b986c807f

Frequently asked questions

npx skillmds add qhjqhj00/longreward-eval