tempreason-eval
Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models — Tan et al. (2023) (arXiv:2306.08952, 2023)
What this evaluates
Evaluates large language models' ability to perform temporal reasoning across three complexity levels: time-time relations (L1), time-event relations (L2), and event-event relations (L3). It specifically probes models' robustness to historical and futuristic time periods, as well as their capacity for month-level intra-year reasoning.
Datasets
- TEMPREASON — total ?; splits: test (-1)
Metrics
EM(primary) — range: [0, 1]- Exact-match accuracy: 1 if the model's predicted answer exactly matches the ground truth string, 0 otherwise.
F1— range: [0, 1]- Token-level F1 score between the predicted temporal expression and the ground truth. Calculated as 2 * (precision * recall) / (precision + recall).
Input / output format
Input: A question requiring temporal reasoning, optionally accompanied by a context passage (CBQA setting) or provided with candidate answers and timestamps (ReasonQA/OBQA settings).
Output: A temporal expression (e.g., year, month, or date span) or a specific answer string corresponding to the question.
Scoring recipe
def compute_metrics(predictions, golds):
em_scores = [1.0 if p == g else 0.0 for p, g in zip(predictions, golds)]
f1_scores = []
for p, g in zip(predictions, golds):
p_tok = set(p.lower().split())
g_tok = set(g.lower().split())
if not p_tok or not g_tok:
f1_scores.append(0.0)
continue
prec = len(p_tok & g_tok) / len(p_tok)
rec = len(p_tok & g_tok) / len(g_tok)
f1_scores.append(2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0.0)
return {'EM': sum(em_scores)/len(em_scores), 'F1': sum(f1_scores)/len(f1_scores)}
Common pitfalls
- Performance varies drastically by question setting (CBQA vs ReasonQA vs OBQA), with CBQA being significantly harder due to lack of context.
- Models exhibit strong temporal bias, performing poorly on pre-1900 and post-2020 years due to pre-training data distribution.
- Intra-year questions requiring month-level reasoning are much harder than inter-year questions, often causing evaluation errors if only year-level matching is used.
- Reasoning shortcuts (e.g., for 'P39 position held' questions) can artificially inflate scores in the CBQA setting.
Evidence (verbatim from paper)
Table 4: Experimental results of each setting in TEMPREASON. Δ F1 refers to the F1 difference between TempT5 and T5-SFT. The reported results are the average scores of three runs.
Citation
@misc{tan2023tempreason,
title={Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models},
author={Tan et al. (2023)},
year={2023},
note={arXiv:2306.08952}
}
- arXiv: 2306.08952