Tempobench Eval

Evaluates large language models' temporal reasoning capabilities by decomposing performance into trace-based (TTE) and causal (TCE) components. It measures how well models handle structured logical specifications with varying complexity, isolating structural factors like horizon depth and information density. Use when the user wants to benchmark on TempoBench, or asks about evaluating this task. Reports exact-match accuracy.

qhjqhj00 6781131 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tempobench-eval commit 67811319fe

Frequently asked questions

npx skillmds add qhjqhj00/tempobench-eval