Mcts Vcb Eval

Evaluates multimodal large language models on fine-grained video captioning by measuring how well generated descriptions capture verified, detailed key points from videos. It probes the model's ability to reason about and describe specific visual details like color, quantity, and position. Use when the user wants to benchmark on MCTS-VCB, or asks about evaluating this task. Reports F1.

qhjqhj00 61802bd 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mcts-vcb-eval commit 61802bdb15

Frequently asked questions

npx skillmds add qhjqhj00/mcts-vcb-eval