mimotable-eval
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning — Li et al. (2024) (arXiv:2412.11711, 2024)
What this evaluates
Evaluates large language models' ability to reason over real-world spreadsheet data, including complex headers, multi-sheet files, and cross-file contexts. It probes capabilities across six meta operations: lookup, edit, calculate, compare, visualize, and reasoning.
Datasets
- MiMoTable — total 1719; splits: test (1719); repo https://github.com/jasonNLP/MiMoTable
Metrics
accuracy(primary) — range: percent- Percentage of questions where the model's generated answer exactly matches the ground truth answer.
Input / output format
Input: A spreadsheet file (or group of files) in .xlsx format, often provided as markdown or raw data, paired with a natural language question.
Output: A natural language answer, a modified spreadsheet file, or a chart, depending on the question type.
Scoring recipe
def compute_accuracy(predictions, gold_answers):
correct = 0
for pred, gold in zip(predictions, gold_answers):
if normalize(pred) == normalize(gold):
correct += 1
return (correct / len(predictions)) * 100
Common pitfalls
- Answers may require code execution or file generation rather than simple text extraction.
- Questions can span multiple sheets or files, requiring cross-context reasoning that standard single-table benchmarks do not test.
- Difficulty grading relies on meta-operation combinations, so models may perform well on simple lookups but fail on reasoning-heavy tasks.
Evidence (verbatim from paper)
Claude-3.5-Sonnet achieve only 77.4% accuracy, indicating significant gaps in handling nuanced table reasoning tasks such as inference and visualization.
Citation
@misc{li2024mimotable,
title={MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning},
author={Li et al. (2024)},
year={2024},
note={arXiv:2412.11711}
}
- arXiv: 2412.11711