Memcollab Eval

Evaluates LLM agents' ability to solve mathematical reasoning and code generation tasks by leveraging a shared, contrastively distilled memory system. It probes cross-agent knowledge transfer, reasoning invariance extraction, and task-aware memory retrieval efficiency. Use when the user wants to benchmark on MATH500, GSM8K, MBPP, HumanEval, or asks about evaluating this task. Reports Accuracy (%).

qhjqhj00 48abedf 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/memcollab-eval commit 48abedf74c

Frequently asked questions

npx skillmds add qhjqhj00/memcollab-eval