# Tempreason Eval

> tempreason-eval

- Skill: `qhjqhj00/tempreason-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/tempreason-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/tempreason-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/tempreason-eval

---


# tempreason-eval

> Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models — Tan et al. (2023) (arXiv:2306.08952, 2023)

## What this evaluates

Evaluates large language models' ability to perform temporal reasoning across three complexity levels: time-time relations (L1), time-event relations (L2), and event-event relations (L3). It specifically probes models' robustness to historical and futuristic time periods, as well as their capacity for month-level intra-year reasoning.

## Datasets

- **TEMPREASON** — total ?; splits: test (-1)

## Metrics

- `EM` **(primary)** — range: [0, 1]
  - Exact-match accuracy: 1 if the model's predicted answer exactly matches the ground truth string, 0 otherwise.
- `F1` — range: [0, 1]
  - Token-level F1 score between the predicted temporal expression and the ground truth. Calculated as 2 * (precision * recall) / (precision + recall).

## Input / output format

**Input**: A question requiring temporal reasoning, optionally accompanied by a context passage (CBQA setting) or provided with candidate answers and timestamps (ReasonQA/OBQA settings).

**Output**: A temporal expression (e.g., year, month, or date span) or a specific answer string corresponding to the question.

## Scoring recipe

```python
def compute_metrics(predictions, golds):
    em_scores = [1.0 if p == g else 0.0 for p, g in zip(predictions, golds)]
    f1_scores = []
    for p, g in zip(predictions, golds):
        p_tok = set(p.lower().split())
        g_tok = set(g.lower().split())
        if not p_tok or not g_tok:
            f1_scores.append(0.0)
            continue
        prec = len(p_tok & g_tok) / len(p_tok)
        rec = len(p_tok & g_tok) / len(g_tok)
        f1_scores.append(2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0.0)
    return {'EM': sum(em_scores)/len(em_scores), 'F1': sum(f1_scores)/len(f1_scores)}
```

## Common pitfalls

- Performance varies drastically by question setting (CBQA vs ReasonQA vs OBQA), with CBQA being significantly harder due to lack of context.
- Models exhibit strong temporal bias, performing poorly on pre-1900 and post-2020 years due to pre-training data distribution.
- Intra-year questions requiring month-level reasoning are much harder than inter-year questions, often causing evaluation errors if only year-level matching is used.
- Reasoning shortcuts (e.g., for 'P39 position held' questions) can artificially inflate scores in the CBQA setting.

## Evidence (verbatim from paper)

> Table 4: Experimental results of each setting in TEMPREASON. Δ F1 refers to the F1 difference between TempT5 and T5-SFT. The reported results are the average scores of three runs.

## Citation

```bibtex
@misc{tan2023tempreason,
  title={Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models},
  author={Tan et al. (2023)},
  year={2023},
  note={arXiv:2306.08952}
}
```

- arXiv: 2306.08952

