Locomo10 Eval

This benchmark evaluates long-term conversational memory systems by testing their ability to retrieve relevant dialogue turns and answer questions over extended, multi-session histories. It probes semantic reasoning, temporal tracking, and adversarial robustness across five distinct question categories. Use when the user wants to benchmark on LoCoMo10, or asks about evaluating this task. Reports F1 score.

qhjqhj00 bd54a94 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/locomo10-eval commit bd54a9460f

Frequently asked questions

npx skillmds add qhjqhj00/locomo10-eval