Refereebench Eval

Evaluates Multimodal Large Language Models (MLLMs) on automatic sports refereeing tasks, probing their ability to detect incidents, classify fouls, apply sport-specific rules, and ground decisions temporally across 11 different sports. Use when the user wants to benchmark on RefereeBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 c1b9fae 2.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/refereebench-eval commit c1b9faefd8

Frequently asked questions

npx skillmds add qhjqhj00/refereebench-eval