Dvbench Eval

Evaluates Vision Large Language Models' ability to understand safety-critical driving videos across a hierarchical taxonomy of 25 abilities, including perception, temporal-spatial reasoning, and risk assessment. Use when the user wants to benchmark on DVBench, or asks about evaluating this task. Reports Top-1 accuracy.

qhjqhj00 bf3874f 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dvbench-eval commit bf3874fbc4

Frequently asked questions

npx skillmds add qhjqhj00/dvbench-eval