Empo2 Memory Augmented LLM Agent

Improve exploration in LLM-based agents through external memory-augmented RL with hybrid on/off-policy training. Agents generate exploration 'tips' (self-reflections) after trajectories, storing them in memory. During rollouts, policy samples between standard execution and memory-conditioned execution. Off-policy updates distill memory-guided behaviors into base policy via reward-guided knowledge distillation. Achieves 128.6% improvement on ScienceWorld and 11.3% on WebShop vs. GRPO.

adu2021 aabf8be 10.1 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/empo2-memory-augmented-llm-agent commit aabf8be948

Frequently asked questions

npx skillmds@latest add adu2021/empo2-memory-augmented-llm-agent