# Tsqa Eval

> tsqa-eval

- Skill: `qhjqhj00/tsqa-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/tsqa-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/tsqa-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/tsqa-eval

---


# tsqa-eval

> Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement — Kong et al. (2025) (arXiv:2503.01875, 2025)

## What this evaluates

Evaluates large language models on time series question answering across five tasks: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. It probes the model's ability to integrate textual context with numerical time series data for both precise numerical prediction and natural language explanation.

## Datasets

- **TSQA** — total 200000; splits: test (250)

## Metrics

- `average MSE` — range: [0, ∞)
  - Mean Squared Error averaged across all forecasting and imputation instances. Calculated as the mean of squared differences between predicted and actual values. Lower values indicate better performance.
- `accuracy` **(primary)** — range: [0, 1]
  - Percentage of correctly predicted labels or answers across anomaly detection, classification, judgment (true-false), and multiple-choice questions. Calculated as correct predictions divided by total instances. Higher values indicate better performance.

## Input / output format

**Input**: Time series data paired with a natural language question/query. Questions and answers are clearly labeled and tokenized.

**Output**: Numerical predictions for forecasting/imputation tasks; class labels for anomaly detection/classification; natural language answers with reasoning for open-ended reasoning (judgment and MCQ) tasks.

## Scoring recipe

```python
def compute_metrics(predictions, gold, tasks):
    results = {}
    for pred, gold_val, task in zip(predictions, gold, tasks):
        if task in ['forecasting', 'imputation']:
            results.setdefault('average MSE', []).append((pred - gold_val) ** 2)
        else:
            results.setdefault('accuracy', []).append(1.0 if pred == gold_val else 0.0)
    return {k: sum(v)/len(v) for k, v in results.items()}
```

## Common pitfalls

- The test set is extremely small (only 50 samples per task), making reported metrics highly susceptible to sampling variance.
- Forecasting tasks use long time series, which the authors explicitly note leads to relatively high MSE values compared to imputation tasks.
- Open-ended reasoning combines multiple-choice and true-false formats but evaluates both solely by accuracy, masking potential differences in question difficulty.

## Evidence (verbatim from paper)

> For evaluation, we randomly selected 50 QA pairs for each task type (or question format). Forecasting and imputation tasks were evaluated using average MSE, while anomaly detection, classification, and open-ended reasoning tasks (including multiple-choice questions (MCQs) and true-false questions (Judgment)) were measured by accuracy. A lower value of MSE ↓ and a higher value of accuracy ↑ indicate better performance.

## Citation

```bibtex
@misc{kong2025timemqa,
  title={Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement},
  author={Kong et al. (2025)},
  year={2025},
  note={arXiv:2503.01875}
}
```

- arXiv: 2503.01875

