Car Bench Eval

Evaluates LLM agents' ability to resolve uncertainty and adhere to safety policies in automotive environments. It probes limit-awareness, consistency across multiple attempts, and robust multi-turn tool use under incomplete or ambiguous user requests. Use when the user wants to benchmark on CAR-bench, or asks about evaluating this task. Reports Passˆ3.

qhjqhj00 4b12ff0 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/car-bench-eval commit 4b12ff0655

Frequently asked questions

npx skillmds add qhjqhj00/car-bench-eval