dlue-eval
DLUE: Benchmarking Document Language Understanding — Xu et al. (2023) (arXiv:2305.09520, 2023)
What this evaluates
Evaluates document language understanding across four core capabilities: classification, structural analysis, information extraction, and transcription. It probes models' ability to handle long documents, complex hierarchical structures, and dispersed knowledge spread across large contexts.
Datasets
- Hyperpartisan — total ?; splits: test (-1)
- ContractNLI — total ?; splits: test (-1)
- ECOM — total ?; splits: test (-1)
- RR — total ?; splits: test (-1)
- GUM — total ?; splits: test (-1)
- LitBank — total ?; splits: test (-1)
- NarrativeQA — total ?; splits: test (-1)
- SummScreen — total ?; splits: test (-1)
- GovReport — total ?; splits: test (-1)
- Qasper — total ?; splits: test (-1)
Metrics
F1(primary) — range: percent- Standard F1 score computed as the harmonic mean of precision and recall. For classification tasks, accuracy is also reported. Values are averaged across three random seed repetitions.
Input / output format
Input: Document text tokenized into sequences. For classification: [CLS] token prepended to the document. For structure analysis: sentence-level sequences with [CLS] tokens at the start of each sentence. For extraction: multi-span question-answering format. For transcription: encoder-decoder input format.
Output: Class label for classification; sequence labels for structure analysis; extracted text spans for information extraction; generated text for transcription.
Scoring recipe
def compute_f1(predictions, gold):
tp = sum(1 for p, g in zip(predictions, gold) if p == g)
fp = sum(1 for p, g in zip(predictions, gold) if p != g and p in gold)
fn = sum(1 for p, g in zip(predictions, gold) if g not in predictions)
precision = tp / (tp + fp) if (tp + fp) > 0 else 0
recall = tp / (tp + fn) if (tp + fn) > 0 else 0
return 2 * precision * recall / (precision + recall) if (precision + recall) > 0 else 0
Common pitfalls
- ContractNLI exhibits label bias related to document length, where longer contracts tend to entail hypotheses, artificially inflating performance on longer documents.
- No single architecture dominates all tasks; performance varies significantly across classification, structure, extraction, and transcription, requiring task-specific model selection.
- Long-range transformers may plateau in performance when document length exceeds input limits, failing to capture dispersed knowledge effectively.
Evidence (verbatim from paper)
human agreement on ECOM was measured at 80.8% F1 (Xu et al., 2022), much higher than our best baseline of 39.1% F1. Likewise, Dasigi et al. (2021) study a subset of Qasper that has multiple annotated answers, and find their overlap to be 60.9% F1, more than double our best baseline.
Citation
@misc{xu2023dlue,
title={DLUE: Benchmarking Document Language Understanding},
author={Xu et al. (2023)},
year={2023},
note={arXiv:2305.09520}
}
- arXiv: 2305.09520