Cleanrl

Single-file deep reinforcement learning implementations (CleanRL). High-quality standalone implementations of PPO, DQN, C51, SAC, DDPG, TD3 with research-friendly features. Each algorithm is a self-contained file with ~300-500 lines. Includes Atari, MuJoCo, Procgen, PettingZoo multi-agent, and JAX variants. Use for RL algorithm reference, rapid prototyping, and understanding implementation details.

mkurman 3767f06 6.2 KB Updated

File contents

mkurman/zorai/tree/main/skills/scientific-skills/cleanrl commit 3767f06e2d

Frequently asked questions

npx skillmds@latest add mkurman/cleanrl