Mem Gallery Eval

This benchmark evaluates multimodal long-term conversational memory in MLLM agents across multi-session dialogues. It probes the agent's ability to extract, adapt, reason over, and manage evolving visual and textual information, including handling temporal dependencies, conflicting updates, and knowledge gaps. Use when the user wants to benchmark on Mem-Gallery, or asks about evaluating this task. Reports answer correctness.

qhjqhj00 e440d53 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mem-gallery-eval commit e440d533a2

Frequently asked questions

npx skillmds add qhjqhj00/mem-gallery-eval