Longgenbench Eval

Evaluates the ability of LLMs to maintain accuracy and logical consistency when generating long-text responses that answer multiple sequential questions from GSM8K or MMLU in a single pass. It specifically probes performance degradation as the number of generated questions increases. Use when the user wants to benchmark on LongGenBench-GSM8K, LongGenBench-MMLU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8adf391 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/longgenbench-eval commit 8adf391e6b

Frequently asked questions

npx skillmds add qhjqhj00/longgenbench-eval