Triviaqa Eval

This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).

qhjqhj00 c8b11b3 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/triviaqa-eval commit c8b11b382f

Frequently asked questions

npx skillmds add qhjqhj00/triviaqa-eval