Muld Eval

Evaluates models' ability to process and extract information from long documents (minimum 10,000 tokens) across multiple NLP tasks including question answering, summarization, classification, and translation. It specifically probes long-context dependency handling and real-world document understanding capabilities. Use when the user wants to benchmark on MuLD Benchmark, NarrativeQA, HotpotQA, OpenSubtitles, or asks about evaluating this task. Reports results.

qhjqhj00 8e487e6 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/muld-eval commit 8e487e6ac3

Frequently asked questions

npx skillmds add qhjqhj00/muld-eval