Using Deep Rl

Use when training, debugging, or selecting a deep-RL algorithm — value-based (DQN/Rainbow/BBF), policy-gradient (PPO/GRPO), actor-critic (SAC/TD3/REDQ), model-based (DreamerV3/TD-MPC2), offline (CQL/IQL/Decision Transformer), multi-agent (MAPPO/IPPO), exploration, reward shaping, or counterfactual credit assignment. Routes to the matching specialist sheet by problem type and algorithm family.

tachyon-beep Updated

File contents

tachyon-beep/skillpacks/tree/main/plugins/yzmir-deep-rl/skills/using-deep-rl commit ab1df8686f

Frequently asked questions

npx skillmds@latest add tachyon-beep/using-deep-rl