severe++-eval
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning — Thoker et al. (2025) (arXiv:2504.05706, 2025)
What this evaluates
Evaluates the generalization and sensitivity of video self-supervised learning models across four factors: domain shift, sample efficiency, action granularity, and task diversity. It probes how well pre-trained representations transfer to diverse downstream datasets, varying finetuning sample sizes, fine-grained actions, and tasks beyond standard action recognition.
Datasets
- UCF-101 — total ?; splits: test (-1)
- NTU-60 — total ?; splits: test (-1)
- FineGym (Gym-99) — total ?; splits: test (-1)
- Something-Something-v2 — total ?; splits: test (-1)
- EPIC-Kitchens-100 — total ?; splits: test (-1)
- Charades — total ?; splits: test (-1)
- AVA — total ?; splits: test (-1)
- ActivityNet — total ?; splits: test (-1)
Metrics
Action Recognition Accuracy(primary) — range: [0, 1]- Standard classification accuracy calculated as the number of correctly predicted labels divided by the total number of samples. For retrieval tasks, mean Average Precision (mAP) is typically used.
Input / output format
Input: Video frames processed through a pre-trained backbone (R(2+1)D-18 for CNNs or ViT-B for transformers) to extract features, followed by a task-dependent linear head for finetuning.
Output: Class label predictions for the corresponding downstream task.
Scoring recipe
def compute_accuracy(predictions, gold_labels):
correct = sum(1 for p, g in zip(predictions, gold_labels) if p == g)
return correct / len(gold_labels)
Common pitfalls
- Assuming pre-training and downstream datasets share the same domain (e.g., all single-camera YouTube videos), ignoring significant domain shift.
- Evaluating only with full downstream data, thereby missing sensitivity to sample efficiency and few-shot transfer.
- Focusing exclusively on action recognition and neglecting task diversity (e.g., video retrieval or temporal reasoning).
- Comparing models trained on different backbones or pre-training regimes without controlling for architectural differences.
Evidence (verbatim from paper)
Most recent works i.e., transformer-based, use action recognition on Kinetics-400 [23] and Something-Something v2 [26] as the main benchmark, and sometimes report additional results on non-standard datasets like AVA [33] or Diving-48 [29]. The problem with these evaluation setups is that the downstream dataset is either the same or shares many similarities with the pre-training dataset... Hence, we believe the current benchmark standard is insufficiently equipped to gain a true understanding of where video self-supervised models are successful, as it cannot show the generalizability or the sensitivity of methods to factors such as domain shift, amount of finetuning data samples, action similarity or task shift.
Citation
@misc{thoker2025severe++,
title={SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning},
author={Thoker et al. (2025)},
year={2025},
note={arXiv:2504.05706}
}
- arXiv: 2504.05706