Agentdrive Mcq Eval

Evaluates large language models' context-sensitive reasoning and decision-making capabilities in autonomous driving scenarios. It probes physics-based calculations, policy compliance, risk interpretation, and maneuver optimization through multiple-choice questions derived from structured driving simulations. Use when the user wants to benchmark on AgentDrive-MCQ, or asks about evaluating this task. Reports accuracy.

qhjqhj00 406e5de 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/agentdrive-mcq-eval commit 406e5de2bd

Frequently asked questions

npx skillmds add qhjqhj00/agentdrive-mcq-eval