mimic-iii-clinical-benchmarks
Attend and Diagnose: Clinical Time Series Analysis using Attention Models — Song et al. (2017) (arXiv:1711.03905, 2017)
What this evaluates
Evaluates clinical time-series models on four critical care prediction tasks: in-hospital mortality, physiologic decompensation, length of stay forecasting, and acute care phenotyping. Tests the model's ability to handle multivariate temporal data, capture long-range dependencies, and perform binary, multi-class, and multi-label classification on real-world ICU records.
Datasets
- MIMIC-III — total 42276; splits: train (-1), val (-1), test (-1)
Metrics
test performance(primary) — range: [0, 1]- Standard classification metrics (e.g., AUC-ROC, accuracy, F1) computed on held-out test sets. The paper refers to this collectively as test performance and optimizes task-specific loss functions during training.
Input / output format
Input: Multivariate clinical time-series data represented as 76-dimensional vectors at each time-step. Time windows vary by task: 24-hour windows for mortality and decompensation, hourly intervals for length of stay, and full ICU stay sequences for phenotyping.
Output: Task-specific predictions: binary classification (mortality/decompensation), 10-class classification (length of stay), or multi-label classification over 25 disease conditions (phenotyping).
Scoring recipe
def compute_test_performance(predictions, gold_labels, task):
if task in ['mortality', 'decompensation']:
return auc(gold_labels, predictions)
elif task == 'los':
return accuracy(gold_labels, predictions)
elif task == 'phenotyping':
return f1_score(gold_labels, predictions, average='micro')
return None
Common pitfalls
- Temporal leakage: ensuring no future data leaks into the 24-hour or hourly prediction windows.
- Class imbalance: mortality (
13%) and decompensation (4.2%) are highly imbalanced, requiring careful thresholding or loss weighting. - Task-specific windowing: decompensation requires step-wise prediction at each time-step, unlike mortality which is episode-level.
Evidence (verbatim from paper)
We illustrate the training behavior and impact of the choice of the attention mask size, number of attention layers and dense interpolation factor on test performance.
Citation
@misc{song2017attend,
title={Attend and Diagnose: Clinical Time Series Analysis using Attention Models},
author={Song et al. (2017)},
year={2017},
note={arXiv:1711.03905}
}
- arXiv: 1711.03905