Scenepilot Bench Eval

Evaluates vision-language models on autonomous driving tasks, including scene understanding, spatial perception, and motion planning. It probes the models' ability to reason about driving scenarios, predict trajectories, and generalize across different geographic regions and traffic conventions. Use when the user wants to benchmark on ScenePilot-Bench, or asks about evaluating this task. Reports Overall Score.

qhjqhj00 deecb66 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/scenepilot-bench-eval commit deecb66fdb

Frequently asked questions

npx skillmds add qhjqhj00/scenepilot-bench-eval