Mia Eval

Evaluates the reasoning, tool-use, and memory-augmented planning capabilities of agents on complex multi-hop QA and visual question answering tasks. It probes how well models leverage episodic memory, test-time learning, and reflection to improve answer accuracy over iterative search trajectories. Use when the user wants to benchmark on FVQA-test, InfoSeek, MMSearch, SimpleVQA, LiveVQA, In-house 1, In-house 2, HotpotQA, 2WikiMultiHopQA, SimpleQA, GAIA (text-only subset), or asks about evaluating this task. Reports accuracy.

qhjqhj00 6ec6171 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mia-eval commit 6ec6171a9f

Frequently asked questions

npx skillmds add qhjqhj00/mia-eval