Benchmark agent memory and RAG systems with MemoryBench
Use MemoryBench to run repeatable conversational memory and RAG benchmarks across providers, datasets, judge models, checkpoints, and structured reports.
Prerequisites
Bun, MemoryBench repository, at least one memory/RAG provider API key, at least one judge model API key, benchmark datasets
Installation
Basic usage or getting-started notes:
🆚 Multi‑provider comparison: run the same benchmark across providers side‑by‑side
📊 Structured reports: export run status, failures, and metrics for analysis
bun install
Extracted from upstream docs: https://raw.githubusercontent.com/supermemoryai/memorybench/HEAD/README.md