Memorybench Eval

This benchmark evaluates how well LLM-based systems retain and utilize both declarative and procedural memory across diverse domains and task formats. It specifically probes continual learning capabilities by measuring performance improvements when systems process explicit and implicit user feedback over multiple interaction sessions. Use when the user wants to benchmark on MemoryBench (Domain & Task Format Partitions), or asks about evaluating this task. Reports LLM-as-Judge score.

qhjqhj00 cb5a6dd 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/memorybench-eval commit cb5a6ddefa

Frequently asked questions

npx skillmds add qhjqhj00/memorybench-eval