Im Promptu Eval

Evaluates an agent's ability to perform in-context compositional reasoning from image prompts by generalizing learned primitive relations to unseen source-target pairs and complex composite tasks. Use when the user wants to benchmark on 3D Shapes, BitMoji Faces, CLEVR Objects, or asks about evaluating this task. Reports MSE.

qhjqhj00 8d1401e 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/im-promptu-eval commit 8d1401e6d2

Frequently asked questions

npx skillmds add qhjqhj00/im-promptu-eval