Loogle Eval

Evaluates the ability of language models to comprehend and reason over long documents (up to 32k+ tokens) by testing short and long dependency tasks, including question answering, cloze completion, and summarization. Use when the user wants to benchmark on LooGLE, or asks about evaluating this task. Reports GPT4_score.

qhjqhj00 2550a32 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/loogle-eval commit 2550a32a65

Frequently asked questions

npx skillmds add qhjqhj00/loogle-eval