# Memoryagentbench Eval

> memoryagentbench-eval

- Skill: `qhjqhj00/memoryagentbench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/memoryagentbench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/memoryagentbench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/memoryagentbench-eval

---


# memoryagentbench-eval

> Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions — Hu et al. (2025) (arXiv:2507.05257, 2025)

## What this evaluates

Evaluates four core memory competencies in LLM agents: accurate retrieval, test-time learning, long-range understanding, and selective forgetting. It transforms long-context datasets into session-based multi-turn interactions to simulate real-world memory accumulation and retrieval.

## Datasets

- **MemoryAgentBench** — total ?; splits: test (-1); repo https://github.com/HUST-AI-HYZ/MemoryAgentBench

## Metrics

- `accuracy` **(primary)** — range: percent
  - Percentage of correctly answered queries against ground-truth labels. Calculated as (number of correct predictions / total number of instances) * 100.

## Input / output format

**Input**: Multi-turn conversational sessions containing long-context documents or prior interaction history, segmented into chunks (typically 512 or 4096 tokens) for retrieval-augmented or long-context processing.

**Output**: Natural language response or structured answer to a specific query posed within the multi-turn session.

## Scoring recipe

```python
correct = 0
for pred, gold in zip(predictions, gold_labels):
    if normalize(pred) == normalize(gold):
        correct += 1
return (correct / len(predictions)) * 100
```

## Common pitfalls

- Chunk size selection significantly biases results: smaller chunks (512) favor RAG agents on retrieval tasks but hurt long-range understanding, while larger chunks (4096) favor long-context models.
- Retrieval top-k is fixed at 10 in main results; increasing it beyond 10 exceeds typical context windows (~40k tokens) and is not evaluated.
- Selective forgetting tasks are extremely difficult for multi-hop scenarios, with most agents achieving ≤7% accuracy, making it a key differentiator from simple retrieval benchmarks.

## Evidence (verbatim from paper)

> The evaluation metrics for all datasets are shown in Table [1], along with more dataset details. ... We observe that all methods fail on the multi-hop situation (with achieving at most 7% accuracy).

## Citation

```bibtex
@misc{hu2025evaluatingmemoryllmagents,
  title={Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions},
  author={Hu et al. (2025)},
  year={2025},
  note={arXiv:2507.05257}
}
```

- arXiv: 2507.05257

