entity6k-eval
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition — Qiu et al. (2024) (arXiv:2403.12339, 2024)
What this evaluates
Evaluates open-domain entity recognition capabilities across four visual grounding and understanding tasks: object detection, zero-shot image classification, image captioning, and dense captioning. It measures how well models can localize, classify, and describe specific real-world entities in images without fine-tuning.
Datasets
- Entity6K — total ?; splits: test (-1)
Metrics
AP (primary) — range: [0, 1]
- Average Precision computed over predicted bounding boxes and ground truth labels, typically averaged across IoU thresholds.
Accuracy — range: [0, 1]
- Standard classification accuracy: the proportion of images correctly assigned to their true class label.
BLEU/ROUGE/METEOR/BertScore — range: [0, 1]
- Standard n-gram overlap (BLEU, ROUGE), word-level alignment (METEOR), and embedding-based similarity (BertScore) between generated captions and ground truth references.
mAP — range: [0, 1]
- Mean Average Precision averaged across pairwise thresholds: IoU thresholds of .3, .4, .5, .6, .7 and METEOR thresholds of 0, .05, .1, .15, .2, .25.
Input / output format
Input: Image (and optional text prompt for classification/captioning tasks).
Output: Object detection: bounding boxes and class labels. Zero-shot classification: predicted class label. Image captioning: generated text description. Dense captioning: bounding boxes and descriptive text for each entity.
Scoring recipe
def compute_metrics(predictions, gold, task):
if task == 'detection':
return compute_AP(predictions, gold, iou_thresholds=[0.5])
elif task == 'classification':
return sum(p == g for p, g in zip(predictions, gold)) / len(gold)
elif task == 'captioning':
return bleu, rouge, meteor, bertscore(predictions, gold)
elif task == 'dense_captioning':
ious = [0.3, 0.4, 0.5, 0.6, 0.7]
meteors = [0.0, 0.05, 0.1, 0.15, 0.2, 0.25]
aps = []
for iou in ious:
for met in meteors:
aps.append(compute_AP(predictions, gold, iou=iou, met_threshold=met))
return mean(aps)
Common pitfalls
- Fine-tuning or training the baseline models, which violates the explicit zero-shot/frozen weight evaluation protocol.
- Using non-standard IoU or METEOR thresholds for dense captioning, as the benchmark requires averaging across the specific ranges (.3-.7 for IoU, .0-.25 for METEOR).
- Ignoring the exact prompt/instruction format required by each baseline model, which can significantly alter zero-shot performance.
Evidence (verbatim from paper)
For object detection, we select Average Precision (AP) as the evaluation metric. For zero-shot image classification, we take the standard accuracy as the evaluation metric. For image captioning, we adopted the BLEU, ROUGE, Meteor, and BertScore as evaluation metrics. For the dense captioning task, we take mean Average Precision (mAP) as the evaluation metric.
Citation
@misc{qiu2024entity6k,
title={Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition},
author={Qiu et al. (2024)},
year={2024},
note={arXiv:2403.12339}
}
1---2name: entity6k-eval3description: entity6k-eval4---56# entity6k-eval78> Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition — Qiu et al. (2024) (arXiv:2403.12339, 2024)910## What this evaluates1112Evaluates open-domain entity recognition capabilities across four visual grounding and understanding tasks: object detection, zero-shot image classification, image captioning, and dense captioning. It measures how well models can localize, classify, and describe specific real-world entities in images without fine-tuning.1314## Datasets1516- **Entity6K** — total ?; splits: test (-1)1718## Metrics1920- `AP` **(primary)** — range: [0, 1]21 - Average Precision computed over predicted bounding boxes and ground truth labels, typically averaged across IoU thresholds.22- `Accuracy` — range: [0, 1]23 - Standard classification accuracy: the proportion of images correctly assigned to their true class label.24- `BLEU/ROUGE/METEOR/BertScore` — range: [0, 1]25 - Standard n-gram overlap (BLEU, ROUGE), word-level alignment (METEOR), and embedding-based similarity (BertScore) between generated captions and ground truth references.26- `mAP` — range: [0, 1]27 - Mean Average Precision averaged across pairwise thresholds: IoU thresholds of .3, .4, .5, .6, .7 and METEOR thresholds of 0, .05, .1, .15, .2, .25.2829## Input / output format3031**Input**: Image (and optional text prompt for classification/captioning tasks).3233**Output**: Object detection: bounding boxes and class labels. Zero-shot classification: predicted class label. Image captioning: generated text description. Dense captioning: bounding boxes and descriptive text for each entity.3435## Scoring recipe3637```python38def compute_metrics(predictions, gold, task):39 if task == 'detection':40 return compute_AP(predictions, gold, iou_thresholds=[0.5])41 elif task == 'classification':42 return sum(p == g for p, g in zip(predictions, gold)) / len(gold)43 elif task == 'captioning':44 return bleu, rouge, meteor, bertscore(predictions, gold)45 elif task == 'dense_captioning':46 ious = [0.3, 0.4, 0.5, 0.6, 0.7]47 meteors = [0.0, 0.05, 0.1, 0.15, 0.2, 0.25]48 aps = []49 for iou in ious:50 for met in meteors:51 aps.append(compute_AP(predictions, gold, iou=iou, met_threshold=met))52 return mean(aps)53```5455## Common pitfalls5657- Fine-tuning or training the baseline models, which violates the explicit zero-shot/frozen weight evaluation protocol.58- Using non-standard IoU or METEOR thresholds for dense captioning, as the benchmark requires averaging across the specific ranges (.3-.7 for IoU, .0-.25 for METEOR).59- Ignoring the exact prompt/instruction format required by each baseline model, which can significantly alter zero-shot performance.6061## Evidence (verbatim from paper)6263> For object detection, we select Average Precision (AP) as the evaluation metric. For zero-shot image classification, we take the standard accuracy as the evaluation metric. For image captioning, we adopted the BLEU, ROUGE, Meteor, and BertScore as evaluation metrics. For the dense captioning task, we take mean Average Precision (mAP) as the evaluation metric.6465## Citation6667```bibtex68@misc{qiu2024entity6k,69 title={Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition},70 author={Qiu et al. (2024)},71 year={2024},72 note={arXiv:2403.12339}73}74```7576- arXiv: 2403.12339