Helmet Long Context Eval

Evaluates a model's ability to retain, process, and reason over extended contexts (8K to 128K tokens) across retrieval-augmented generation (RAG) and long-range question answering (LongQA) tasks. It probes robustness to noise, multi-hop reasoning, and memorization in long-context settings. Use when the user wants to benchmark on HELMET, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4d4cb1e 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/helmet-long-context-eval commit 4d4cb1e214

Frequently asked questions

npx skillmds add qhjqhj00/helmet-long-context-eval