Task Me Anything Eval

Evaluates the visual perceptual capabilities of large multimodal language models across object recognition, attribute recognition, spatial and temporal reasoning, and action recognition using programmatically generated image and video question-answering tasks. Use when the user wants to benchmark on Task-Me-Anything, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7689d21 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/task-me-anything-eval commit 7689d2127c

Frequently asked questions

npx skillmds add qhjqhj00/task-me-anything-eval