Long Context Retrieval Eval

Probes a model's ability to accurately retrieve hidden, specific information (needles) embedded within extremely long multimodal sequences (text, video, audio) and assesses its predictive stability over millions of tokens. Use when the user wants to benchmark on Paul Graham Essays (Synthetic), AlphaGo Documentary, VoxPopuli, or asks about evaluating this task. Reports recall.

qhjqhj00 946c71f 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/long-context-retrieval-eval commit 946c71f3b1

Frequently asked questions

npx skillmds add qhjqhj00/long-context-retrieval-eval