Drivebench Eval

Evaluates the reliability, visual grounding, and corruption resilience of vision-language models in autonomous driving. It probes whether models genuinely interpret degraded visual inputs or rely on textual priors and hallucinated reasoning when visual cues are missing or corrupted. Use when the user wants to benchmark on DriveBench, or asks about evaluating this task. Reports GPT score.

qhjqhj00 93af9d1 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/drivebench-eval commit 93af9d16eb

Frequently asked questions

npx skillmds add qhjqhj00/drivebench-eval