# Treb Eval

> treb-eval

- Skill: `qhjqhj00/treb-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/treb-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/treb-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/treb-eval

---


# treb-eval

> TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models — Li et al. (2025) (arXiv:2506.18421, 2025)

## What this evaluates

Evaluates large language models' table reasoning capabilities across six core skills: natural language understanding, table understanding, basic/computational operations, data analysis, and advanced data analysis. It tests both direct textual reasoning and code-based reasoning (Program-of-Thought and Interleaved CoT) for extracting and processing tabular information.

## Datasets

- **TReB** — total ?; splits: test (-1); repo https://github.com/JT-LM/jiutian-treb

## Metrics

- `ROUGE-L` **(primary)** — range: [0, 100]
  - Measures the longest common subsequence overlap between the generated answer and the ground truth answer, reported as a percentage.
- `LLM-as-a-judge` — range: [0, 100]
  - Uses Qwen2-72B-Instruct to evaluate semantic similarity and answer accuracy between the model's output and the ground truth.
- `Accuracy` — range: [0, 1]
  - Exact match score calculated only for tasks where the ground truth answer is a single numerical value.

## Input / output format

**Input**: For TCoT: table content in Markdown/HTML format plus a question. For PoT/ICoT: plaintext table content is not provided; the model receives a prompt requiring it to generate code to read and process the table.

**Output**: For TCoT: a direct textual answer. For PoT/ICoT: executable code to extract/process information, from which the final answer is parsed. Strict formatting constraints are applied to isolate the final answer.

## Scoring recipe

```python
if mode == 'TCoT':
    score = rouge_l(generated_text, ground_truth)
elif mode in ['PoT', 'ICoT']:
    try:
        execution_result = execute_code(generated_code)
        score = rouge_l(execution_result, ground_truth)
    except ExecutionError:
        score = 0.0
if is_single_numerical_answer(ground_truth):
    score = 1.0 if generated_text == ground_truth else 0.0
return score
```

## Common pitfalls

- PoT and ICoT tasks automatically receive a score of zero if the generated code fails to execute.
- Different inference modes are task-specific; e.g., NLU tasks exclusively use TCoT, so applying PoT/ICoT to them is invalid.
- Strict output formatting constraints are enforced to extract final answers; models must avoid irrelevant text that could break parsing.

## Evidence (verbatim from paper)

> In the following experiments, we primarily use ROUGE-L*(Lin, [2004])* and LLM-as-a-judge*(Zheng et al., [2023])* to evaluate model performance across tasks. ROUGE-L assesses the textual similarity between the student answer and the ground truth answer, while LLM-as-a-judge evaluates semantic similarity and answer accuracy.

## Citation

```bibtex
@misc{li2025treb,
  title={TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models},
  author={Li et al. (2025)},
  year={2025},
  note={arXiv:2506.18421}
}
```

- arXiv: 2506.18421

