ilsvrc-eval
ImageNet Large Scale Visual Recognition Challenge — Olga Russakovsky et al. (arXiv:1409.0575, 2014)
What this evaluates
Evaluates large-scale visual recognition capabilities across three core tasks: image classification, single-object localization, and object detection. It probes a model's ability to categorize, localize, and detect objects across 1,000 diverse categories using a dataset of approximately 1 million images.
Datasets
- ILSVRC — total 1000000; splits: train (-1), val (-1), test (-1)
Metrics
classification error(primary) — range: percent- Computed as 1 minus the top-k accuracy (typically top-5 in early years, top-1 later). Measures the fraction of images where the ground-truth class is not among the top predicted classes.
mean average precision— range: [0, 1]- The average of the area under the precision-recall curve computed per object class, then averaged across all 200 categories in the detection task.
Input / output format
Input: RGB images of varying resolutions, with ground-truth annotations including class labels, bounding box coordinates, and object presence indicators.
Output: For classification: a ranked list of predicted class labels with confidence scores. For localization/detection: a set of predicted bounding boxes with associated class labels and confidence scores.
Scoring recipe
def compute_classification_error(predictions, gold_labels):
correct = sum(1 for p, g in zip(predictions, gold_labels) if g in p[:5])
return 1.0 - (correct / len(gold_labels))
def compute_mAP(predictions, gold_boxes, classes):
aps = []
for cls in classes:
precisions, recalls = compute_pr_curve(predictions, gold_boxes, cls)
aps.append(trapezoidal_rule(precisions, recalls))
return sum(aps) / len(aps)
Common pitfalls
- Starting in 2014, teams were allowed to use external training data, creating separate 'provided data' and 'external data' tracks that must not be mixed when reporting results.
- Single-object localization requires predicting exactly one bounding box per image, whereas object detection requires predicting all instances of multiple classes per image.
- Early ILSVRC years primarily reported top-5 classification error, while later years shifted focus to top-1 error, requiring careful year-specific metric alignment.
Evidence (verbatim from paper)
The ILSVRC dataset and the competition has allowed significant algorithmic advances in large-scale image recognition and retrieval. ... Significant progress has been made in just one year: image classification error was almost halved since ILSVRC2013 and object detection mean average precision almost doubled compared to ILSVRC2013.
Citation
@misc{russakovsky2014ilsvrc,
title={ImageNet Large Scale Visual Recognition Challenge},
author={Olga Russakovsky et al.},
year={2014},
note={arXiv:1409.0575}
}
- arXiv: 1409.0575