# Ilsvrc Eval

> ilsvrc-eval

- Skill: `qhjqhj00/ilsvrc-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/ilsvrc-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/ilsvrc-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/ilsvrc-eval

---


# ilsvrc-eval

> ImageNet Large Scale Visual Recognition Challenge — Olga Russakovsky et al. (arXiv:1409.0575, 2014)

## What this evaluates

Evaluates large-scale visual recognition capabilities across three core tasks: image classification, single-object localization, and object detection. It probes a model's ability to categorize, localize, and detect objects across 1,000 diverse categories using a dataset of approximately 1 million images.

## Datasets

- **ILSVRC** — total 1000000; splits: train (-1), val (-1), test (-1)

## Metrics

- `classification error` **(primary)** — range: percent
  - Computed as 1 minus the top-k accuracy (typically top-5 in early years, top-1 later). Measures the fraction of images where the ground-truth class is not among the top predicted classes.
- `mean average precision` — range: [0, 1]
  - The average of the area under the precision-recall curve computed per object class, then averaged across all 200 categories in the detection task.

## Input / output format

**Input**: RGB images of varying resolutions, with ground-truth annotations including class labels, bounding box coordinates, and object presence indicators.

**Output**: For classification: a ranked list of predicted class labels with confidence scores. For localization/detection: a set of predicted bounding boxes with associated class labels and confidence scores.

## Scoring recipe

```python
def compute_classification_error(predictions, gold_labels):
    correct = sum(1 for p, g in zip(predictions, gold_labels) if g in p[:5])
    return 1.0 - (correct / len(gold_labels))

def compute_mAP(predictions, gold_boxes, classes):
    aps = []
    for cls in classes:
        precisions, recalls = compute_pr_curve(predictions, gold_boxes, cls)
        aps.append(trapezoidal_rule(precisions, recalls))
    return sum(aps) / len(aps)
```

## Common pitfalls

- Starting in 2014, teams were allowed to use external training data, creating separate 'provided data' and 'external data' tracks that must not be mixed when reporting results.
- Single-object localization requires predicting exactly one bounding box per image, whereas object detection requires predicting all instances of multiple classes per image.
- Early ILSVRC years primarily reported top-5 classification error, while later years shifted focus to top-1 error, requiring careful year-specific metric alignment.

## Evidence (verbatim from paper)

> The ILSVRC dataset and the competition has allowed significant algorithmic advances in large-scale image recognition and retrieval. ... Significant progress has been made in just one year: image classification error was almost halved since ILSVRC2013 and object detection mean average precision almost doubled compared to ILSVRC2013.

## Citation

```bibtex
@misc{russakovsky2014ilsvrc,
  title={ImageNet Large Scale Visual Recognition Challenge},
  author={Olga Russakovsky et al.},
  year={2014},
  note={arXiv:1409.0575}
}
```

- arXiv: 1409.0575

