mosaicdoc-eval
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding — Chen et al. (2025) (arXiv:2511.09919, 2025)
What this evaluates
Probes Visually Rich Document Understanding (VRDU) capabilities across three tasks: document VQA, page-level OCR, and reading order prediction. It specifically tests models' ability to comprehend dense, bilingual, non-Manhattan layouts, perform multi-span reasoning, and maintain global layout coherence without token reduction artifacts.
Datasets
- MosaicDoc — total 72000; splits: test (-1); repo https://github.com/DOCLAB-SCUT/MosaicDoc
Metrics
ANLSL(primary) — range: [0, 1]- Average Normalized Levenshtein Similarity for List. Computes the normalized Levenshtein similarity between predicted and ground-truth answer lists, averaged across the dataset.
CRR— range: percent- Character Recognition Rate. Measures character-level accuracy by comparing the model's full page transcription against the ground truth.
OCRR— range: percent- Output-based Character Recognition Rate. Normalizes character accuracy by the length of the model's generated output to measure precision within the prediction.
Micro-F1— range: [0, 1]- Micro-averaged F1 score computed on text line sequences to evaluate reading order prediction correctness by matching predicted blocks to the ground truth sequence.
Input / output format
Input: Image of a visually rich document (magazine or newspaper) paired with a question (for VQA), or image alone (for OCR and reading order prediction).
Output: Text string (answer for VQA, full page transcription for OCR, ordered list of text line blocks for reading order prediction).
Scoring recipe
def compute_anlsl(pred_answers, gold_answers):
sims = [levenshtein_similarity(p, g) for p, g in zip(pred_answers, gold_answers)]
return sum(sims) / len(sims)
def compute_crr(pred_text, gold_text):
return char_accuracy(pred_text, gold_text)
def compute_ocrr(pred_text, gold_text):
return char_accuracy(pred_text, gold_text) / len(pred_text)
def compute_micro_f1(pred_order, gold_order):
tp = sum(1 for p, g in zip(pred_order, gold_order) if p == g)
fp = len(pred_order) - tp
fn = len(gold_order) - tp
return 2 * tp / (2 * tp + fp + fn) if (2 * tp + fp + fn) > 0 else 0.0
Common pitfalls
- Expert VLMs using token reduction/merging strategies lose critical semantic information on dense, information-rich layouts.
- Models frequently exhibit repetitive output generation until hitting token limits, which artificially deflates CRR/OCRR scores.
- Reading order prediction often yields high precision but low recall, as models correctly order local column fragments but fail to capture global multi-column or non-Manhattan sequencing.
Evidence (verbatim from paper)
We evaluate DocVQA performance using the Average Normalized Levenshtein Similarity for List (ANLSL) metric. To assess raw text recognition capability, we evaluate page-level Character Recognition Rate (OCR) using CRR and Output-based Character Recognition Rate (OCRR). We evaluate ROP using the Micro-F1 score on text line sequences.
Citation
@misc{chen2025mosaicdoc,
title={MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding},
author={Chen et al. (2025)},
year={2025},
note={arXiv:2511.09919}
}
- arXiv: 2511.09919