Just In Time Reinforcement Learning Continual

Implement JitRL-style continual learning for LLM agents: training-free policy optimization via experience memory, advantage estimation, and logit modulation. Use when asked to 'add experience memory to an agent', 'implement continual learning without fine-tuning', 'build a JitRL agent', 'optimize agent actions from past trajectories', 'add non-parametric RL to an LLM pipeline', or 'make my agent learn from its mistakes at inference time'.

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/just-in-time-reinforcement-learning-continual commit 4a20ef4221

Frequently asked questions

npx skillmds@latest add ndpvt-web/just-in-time-reinforcement-learning-continual