Ad2 Bench Eval

Evaluates multimodal large language models on autonomous driving tasks under adverse weather and complex scenes. It probes base and advanced visual perception, relational understanding, event reasoning, and the coherence of hierarchical chain-of-thought reasoning. Use when the user wants to benchmark on AD^2-Bench, or asks about evaluating this task. Reports Avg-S.

qhjqhj00 de6032e 2.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ad2-bench-eval commit de6032e90e

Frequently asked questions

npx skillmds add qhjqhj00/ad2-bench-eval