Cctvbench Eval

Probes multimodal LLMs' ability to answer binary questions about traffic videos while maintaining logical consistency across counterfactual video-question pairs. It specifically diagnoses failure modes like positive omission, negative hallucination, and mutual-exclusivity violations by enforcing a strict quadruple-level decision rule. Use when the user wants to benchmark on CCTVBench, or asks about evaluating this task. Reports QuadAcc.

qhjqhj00 d07c4c1 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cctvbench-eval commit d07c4c1722

Frequently asked questions

npx skillmds add qhjqhj00/cctvbench-eval