convomem-eval
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG — Pakhomov et al. (2025) (arXiv:2511.10523, 2025)
What this evaluates
Evaluates conversational memory capabilities across six dimensions: recalling user facts, tracking assistant statements, abstaining when information is missing, inferring preferences, handling changing facts, and making implicit connections. It specifically tests the ability to synthesize evidence distributed across multiple conversation turns.
Datasets
- ConvoMem — total 75336; splits: test (75336); repo https://github.com/SalesforceAIResearch/ConvoMem
Metrics
accuracy(primary) — range: percent- Percentage of correctly answered questions out of the total test set. For rubric-based categories (Preferences, Implicit Connections), responses are evaluated against predefined criteria to determine correctness.
Input / output format
Input: A conversational context consisting of 1 to 6 messages containing scattered evidence, followed by a question requiring memory retrieval or synthesis.
Output: A natural language response answering the question. For abstention cases, the model should explicitly state that the information is unavailable.
Scoring recipe
correct = 0
total = len(predictions)
for pred, gold in zip(predictions, golds):
if is_correct(pred, gold): # Exact match or rubric-based evaluation
correct += 1
return (correct / total) * 100
Common pitfalls
- Models may exploit stylistic patterns if evidence and filler conversations are generated by different pipelines, leading to inflated scores without genuine memory retrieval.
- Multi-message evidence requires synthesizing information scattered across multiple turns; systems relying on single-turn retrieval will fail on 60% of test cases.
- Abstention cases deliberately omit answers; models prone to hallucination will fabricate plausible but incorrect responses instead of admitting missing information.
Evidence (verbatim from paper)
A large-scale benchmark of 75,336 question-answer pairs evaluates conversational memory across user facts, preferences, temporal changes, and implicit connections, revealing that simple full-context approaches achieve 70–82% accuracy on multi-message evidence—outperforming RAG-based systems (30–45%) in early conversations.
Citation
@misc{pakhomov2025convomem,
title={Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG},
author={Pakhomov et al. (2025)},
year={2025},
note={arXiv:2511.10523}
}
- arXiv: 2511.10523