Traveler Temporal Reasoning Eval

This benchmark evaluates large language models' ability to perform event-temporal reasoning by resolving explicit, implicit, and vague temporal references across synthetic household event chains. It systematically probes how model performance degrades with increasing event set length and varying levels of temporal explicitness. Use when the user wants to benchmark on TRAVELER, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f5e3f2f 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/traveler-temporal-reasoning-eval commit f5e3f2f6bb

Frequently asked questions

npx skillmds add qhjqhj00/traveler-temporal-reasoning-eval