Stable Baselines3 Algorithm Reference
This document provides detailed characteristics of all RL algorithms in Stable Baselines3 to help select the right algorithm for specific tasks.
Algorithm Comparison Table
| Algorithm |
Type |
Action Space |
Sample Efficiency |
Training Speed |
Use Case |
| PPO |
On-Policy |
All |
Medium |
Fast |
General-purpose, stable |
| A2C |
On-Policy |
All |
Low |
Very Fast |
Quick prototyping, multiprocessing |
| SAC |
Off-Policy |
Continuous |
High |
Medium |
Continuous control, sample-efficient |
| TD3 |
Off-Policy |
Continuous |
High |
Medium |
Continuous control, deterministic |
| DDPG |
Off-Policy |
Continuous |
High |
Medium |
Continuous control (use TD3 instead) |
| DQN |
Off-Policy |
Discrete |
Medium |
Medium |
Discrete actions, Atari games |
| HER |
Off-Policy |
All |
Very High |
Medium |
Goal-conditioned tasks |
| RecurrentPPO |
On-Policy |
All |
Medium |
Slow |
Partial observability (POMDP) |
Detailed Algorithm Characteristics
PPO (Proximal Policy Optimization)
Overview: General-purpose on-policy algorithm with good performance across many tasks.
Strengths:
- Stable and reliable training
- Works with all action space types (Discrete, Box, MultiDiscrete, MultiBinary)
- Good balance between sample efficiency and training speed
- Excellent for multiprocessing with vectorized environments
- Easy to tune
Weaknesses:
- Less sample-efficient than off-policy methods
- Requires many environment interactions
Best For:
- General-purpose RL tasks
- When stability is important
- When you have cheap environment simulations
- Tasks with continuous or discrete actions
Hyperparameter Guidance:
n_steps: 2048-4096 for continuous, 128-256 for Atari
learning_rate: 3e-4 is a good default
n_epochs: 10 for continuous, 4 for Atari
batch_size: 64
gamma: 0.99 (0.995-0.999 for long episodes)
A2C (Advantage Actor-Critic)
Overview: Synchronous variant of A3C, simpler than PPO but less stable.
Strengths:
- Very fast training (simpler than PPO)
- Works with all action space types
- Good for quick prototyping
- Memory efficient
Weaknesses:
- Less stable than PPO
- Requires careful hyperparameter tuning
- Lower sample efficiency
Best For:
- Quick experimentation
- When training speed is critical
- Simple environments
Hyperparameter Guidance:
n_steps: 5-256 depending on task
learning_rate: 7e-4
gamma: 0.99
SAC (Soft Actor-Critic)
Overview: Off-policy algorithm with entropy regularization, state-of-the-art for continuous control.
Strengths:
- Excellent sample efficiency
- Very stable training
- Automatic entropy tuning
- Good exploration through stochastic policy
- State-of-the-art for robotics
Weaknesses:
- Only supports continuous action spaces (Box)
- Slower wall-clock time than on-policy methods
- More complex hyperparameters
Best For:
- Continuous control (robotics, physics simulations)
- When sample efficiency is critical
- Expensive environment simulations
- Tasks requiring good exploration
Hyperparameter Guidance:
learning_rate: 3e-4
buffer_size: 1M for most tasks
learning_starts: 10000
batch_size: 256
tau: 0.005 (target network update rate)
train_freq: 1 with gradient_steps=-1 for best performance
TD3 (Twin Delayed DDPG)
Overview: Improved DDPG with double Q-learning and delayed policy updates.
Strengths:
- High sample efficiency
- Deterministic policy (good for deployment)
- More stable than DDPG
- Good for continuous control
Weaknesses:
- Only supports continuous action spaces (Box)
- Less exploration than SAC
- Requires careful tuning
Best For:
- Continuous control tasks
- When deterministic policies are preferred
- Sample-efficient learning
Hyperparameter Guidance:
learning_rate: 1e-3
buffer_size: 1M
learning_starts: 10000
batch_size: 100
policy_delay: 2 (update policy every 2 critic updates)
DDPG (Deep Deterministic Policy Gradient)
Overview: Early off-policy continuous control algorithm.
Strengths:
- Continuous action space support
- Off-policy learning
Weaknesses:
- Less stable than TD3 or SAC
- Sensitive to hyperparameters
- Generally outperformed by TD3
Best For:
- Legacy compatibility
- Recommendation: Use TD3 instead for new projects
DQN (Deep Q-Network)
Overview: Classic off-policy algorithm for discrete action spaces.
Strengths:
- Sample-efficient for discrete actions
- Experience replay enables reuse of past data
- Proven success on Atari games
Weaknesses:
- Only supports discrete action spaces
- Can be unstable without proper tuning
- Overestimation bias
Best For:
- Discrete action tasks
- Atari games and similar environments
- When sample efficiency matters
Hyperparameter Guidance:
learning_rate: 1e-4
buffer_size: 100K-1M depending on task
learning_starts: 50000 for Atari
batch_size: 32
exploration_fraction: 0.1
exploration_final_eps: 0.05
Variants:
- QR-DQN: Distributional RL version for better value estimates
- Maskable DQN: For environments with action masking
HER (Hindsight Experience Replay)
Overview: Not a standalone algorithm but a replay buffer strategy for goal-conditioned tasks.
Strengths:
- Dramatically improves learning in sparse reward settings
- Learns from failures by relabeling goals
- Works with any off-policy algorithm (SAC, TD3, DQN)
Weaknesses:
- Only for goal-conditioned environments
- Requires specific observation structure (Dict with "observation", "achieved_goal", "desired_goal")
Best For:
- Goal-conditioned tasks (robotics manipulation, navigation)
- Sparse reward environments
- Tasks where goal is clear but reward is binary
Usage:
from stable_baselines3 import SAC, HerReplayBuffer
model = SAC(
"MultiInputPolicy",
env,
replay_buffer_class=HerReplayBuffer,
replay_buffer_kwargs=dict(
n_sampled_goal=4,
goal_selection_strategy="future", # or "episode", "final"
),
)
RecurrentPPO
Overview: PPO with LSTM policy for handling partial observability.
Strengths:
- Handles partial observability (POMDP)
- Can learn temporal dependencies
- Good for memory-required tasks
Weaknesses:
- Slower training than standard PPO
- More complex to tune
- Requires sequential data
Best For:
- Partially observable environments
- Tasks requiring memory (e.g., navigation without full map)
- Time-series problems
Algorithm Selection Guide
Decision Tree
What is your action space?
- Continuous (Box) → Consider PPO, SAC, or TD3
- Discrete → Consider PPO, A2C, or DQN
- MultiDiscrete/MultiBinary → Use PPO or A2C
Is sample efficiency critical?
- Yes (expensive simulations) → Use off-policy: SAC, TD3, DQN, or HER
- No (cheap simulations) → Use on-policy: PPO, A2C
Do you need fast wall-clock training?
- Yes → Use PPO or A2C with vectorized environments
- No → Any algorithm works
Is the task goal-conditioned with sparse rewards?
- Yes → Use HER with SAC or TD3
- No → Continue with standard algorithms
Is the environment partially observable?
- Yes → Use RecurrentPPO
- No → Use standard algorithms
Quick Recommendations
- Starting out / General tasks: PPO
- Continuous control / Robotics: SAC
- Discrete actions / Atari: DQN or PPO
- Goal-conditioned / Sparse rewards: SAC + HER
- Fast prototyping: A2C
- Sample efficiency critical: SAC, TD3, or DQN
- Partial observability: RecurrentPPO
Training Configuration Tips
For On-Policy Algorithms (PPO, A2C)
# Use vectorized environments for speed
env = make_vec_env(env_id, n_envs=8, vec_env_cls=SubprocVecEnv)
model = PPO(
"MlpPolicy",
env,
n_steps=2048, # Collect this many steps per environment before update
batch_size=64,
n_epochs=10,
learning_rate=3e-4,
gamma=0.99,
)
For Off-Policy Algorithms (SAC, TD3, DQN)
# Fewer environments, but use gradient_steps=-1 for efficiency
env = make_vec_env(env_id, n_envs=4)
model = SAC(
"MlpPolicy",
env,
buffer_size=1_000_000,
learning_starts=10000,
batch_size=256,
train_freq=1,
gradient_steps=-1, # Do 1 gradient step per env step (4 with 4 envs)
learning_rate=3e-4,
)
Common Pitfalls
- Using DQN with continuous actions - DQN only works with discrete actions
- Not using vectorized environments with PPO/A2C - Wastes potential speedup
- Using too few environments - On-policy methods need many samples
- Using too large replay buffer - Can cause memory issues
- Not tuning learning rate - Critical for stable training
- Ignoring reward scaling - Normalize rewards for better learning
- Wrong policy type - Use "CnnPolicy" for images, "MultiInputPolicy" for dict observations
Performance Benchmarks
Approximate expected performance (mean reward) on common benchmarks:
Continuous Control (MuJoCo)
- HalfCheetah-v3: PPO ~1800, SAC ~12000, TD3 ~9500
- Hopper-v3: PPO ~2500, SAC ~3600, TD3 ~3600
- Walker2d-v3: PPO ~3000, SAC ~5500, TD3 ~5000
Discrete Control (Atari)
- Breakout: PPO ~400, DQN ~300
- Pong: PPO ~20, DQN ~20
- Space Invaders: PPO ~1000, DQN ~800
Note: Performance varies significantly with hyperparameters and training time.
Additional Resources
- RL Baselines3 Zoo: Collection of pre-trained agents and hyperparameters: https://github.com/DLR-RM/rl-baselines3-zoo
- Hyperparameter Tuning: Use Optuna for systematic tuning
- Custom Policies: Extend base policies for custom network architectures
- Contribution Repo: SB3-Contrib for experimental algorithms (QR-DQN, TQC, etc.)
1---2name: stable-baselines3-algorithm-reference3description: This document provides detailed characteristics of all RL algorithms in Stable Baselines3 to help select the right algorithm for specific tasks.4---5# Stable Baselines3 Algorithm Reference67This document provides detailed characteristics of all RL algorithms in Stable Baselines3 to help select the right algorithm for specific tasks.89## Algorithm Comparison Table1011| Algorithm | Type | Action Space | Sample Efficiency | Training Speed | Use Case |12|-----------|------|--------------|-------------------|----------------|----------|13| **PPO** | On-Policy | All | Medium | Fast | General-purpose, stable |14| **A2C** | On-Policy | All | Low | Very Fast | Quick prototyping, multiprocessing |15| **SAC** | Off-Policy | Continuous | High | Medium | Continuous control, sample-efficient |16| **TD3** | Off-Policy | Continuous | High | Medium | Continuous control, deterministic |17| **DDPG** | Off-Policy | Continuous | High | Medium | Continuous control (use TD3 instead) |18| **DQN** | Off-Policy | Discrete | Medium | Medium | Discrete actions, Atari games |19| **HER** | Off-Policy | All | Very High | Medium | Goal-conditioned tasks |20| **RecurrentPPO** | On-Policy | All | Medium | Slow | Partial observability (POMDP) |2122## Detailed Algorithm Characteristics2324### PPO (Proximal Policy Optimization)2526**Overview:** General-purpose on-policy algorithm with good performance across many tasks.2728**Strengths:**29- Stable and reliable training30- Works with all action space types (Discrete, Box, MultiDiscrete, MultiBinary)31- Good balance between sample efficiency and training speed32- Excellent for multiprocessing with vectorized environments33- Easy to tune3435**Weaknesses:**36- Less sample-efficient than off-policy methods37- Requires many environment interactions3839**Best For:**40- General-purpose RL tasks41- When stability is important42- When you have cheap environment simulations43- Tasks with continuous or discrete actions4445**Hyperparameter Guidance:**46- `n_steps`: 2048-4096 for continuous, 128-256 for Atari47- `learning_rate`: 3e-4 is a good default48- `n_epochs`: 10 for continuous, 4 for Atari49- `batch_size`: 6450- `gamma`: 0.99 (0.995-0.999 for long episodes)5152### A2C (Advantage Actor-Critic)5354**Overview:** Synchronous variant of A3C, simpler than PPO but less stable.5556**Strengths:**57- Very fast training (simpler than PPO)58- Works with all action space types59- Good for quick prototyping60- Memory efficient6162**Weaknesses:**63- Less stable than PPO64- Requires careful hyperparameter tuning65- Lower sample efficiency6667**Best For:**68- Quick experimentation69- When training speed is critical70- Simple environments7172**Hyperparameter Guidance:**73- `n_steps`: 5-256 depending on task74- `learning_rate`: 7e-475- `gamma`: 0.997677### SAC (Soft Actor-Critic)7879**Overview:** Off-policy algorithm with entropy regularization, state-of-the-art for continuous control.8081**Strengths:**82- Excellent sample efficiency83- Very stable training84- Automatic entropy tuning85- Good exploration through stochastic policy86- State-of-the-art for robotics8788**Weaknesses:**89- Only supports continuous action spaces (Box)90- Slower wall-clock time than on-policy methods91- More complex hyperparameters9293**Best For:**94- Continuous control (robotics, physics simulations)95- When sample efficiency is critical96- Expensive environment simulations97- Tasks requiring good exploration9899**Hyperparameter Guidance:**100- `learning_rate`: 3e-4101- `buffer_size`: 1M for most tasks102- `learning_starts`: 10000103- `batch_size`: 256104- `tau`: 0.005 (target network update rate)105- `train_freq`: 1 with `gradient_steps=-1` for best performance106107### TD3 (Twin Delayed DDPG)108109**Overview:** Improved DDPG with double Q-learning and delayed policy updates.110111**Strengths:**112- High sample efficiency113- Deterministic policy (good for deployment)114- More stable than DDPG115- Good for continuous control116117**Weaknesses:**118- Only supports continuous action spaces (Box)119- Less exploration than SAC120- Requires careful tuning121122**Best For:**123- Continuous control tasks124- When deterministic policies are preferred125- Sample-efficient learning126127**Hyperparameter Guidance:**128- `learning_rate`: 1e-3129- `buffer_size`: 1M130- `learning_starts`: 10000131- `batch_size`: 100132- `policy_delay`: 2 (update policy every 2 critic updates)133134### DDPG (Deep Deterministic Policy Gradient)135136**Overview:** Early off-policy continuous control algorithm.137138**Strengths:**139- Continuous action space support140- Off-policy learning141142**Weaknesses:**143- Less stable than TD3 or SAC144- Sensitive to hyperparameters145- Generally outperformed by TD3146147**Best For:**148- Legacy compatibility149- **Recommendation:** Use TD3 instead for new projects150151### DQN (Deep Q-Network)152153**Overview:** Classic off-policy algorithm for discrete action spaces.154155**Strengths:**156- Sample-efficient for discrete actions157- Experience replay enables reuse of past data158- Proven success on Atari games159160**Weaknesses:**161- Only supports discrete action spaces162- Can be unstable without proper tuning163- Overestimation bias164165**Best For:**166- Discrete action tasks167- Atari games and similar environments168- When sample efficiency matters169170**Hyperparameter Guidance:**171- `learning_rate`: 1e-4172- `buffer_size`: 100K-1M depending on task173- `learning_starts`: 50000 for Atari174- `batch_size`: 32175- `exploration_fraction`: 0.1176- `exploration_final_eps`: 0.05177178**Variants:**179- **QR-DQN**: Distributional RL version for better value estimates180- **Maskable DQN**: For environments with action masking181182### HER (Hindsight Experience Replay)183184**Overview:** Not a standalone algorithm but a replay buffer strategy for goal-conditioned tasks.185186**Strengths:**187- Dramatically improves learning in sparse reward settings188- Learns from failures by relabeling goals189- Works with any off-policy algorithm (SAC, TD3, DQN)190191**Weaknesses:**192- Only for goal-conditioned environments193- Requires specific observation structure (Dict with "observation", "achieved_goal", "desired_goal")194195**Best For:**196- Goal-conditioned tasks (robotics manipulation, navigation)197- Sparse reward environments198- Tasks where goal is clear but reward is binary199200**Usage:**201```python202from stable_baselines3 import SAC, HerReplayBuffer203204model = SAC(205 "MultiInputPolicy",206 env,207 replay_buffer_class=HerReplayBuffer,208 replay_buffer_kwargs=dict(209 n_sampled_goal=4,210 goal_selection_strategy="future", # or "episode", "final"211 ),212)213```214215### RecurrentPPO216217**Overview:** PPO with LSTM policy for handling partial observability.218219**Strengths:**220- Handles partial observability (POMDP)221- Can learn temporal dependencies222- Good for memory-required tasks223224**Weaknesses:**225- Slower training than standard PPO226- More complex to tune227- Requires sequential data228229**Best For:**230- Partially observable environments231- Tasks requiring memory (e.g., navigation without full map)232- Time-series problems233234## Algorithm Selection Guide235236### Decision Tree2372381. **What is your action space?**239 - **Continuous (Box)** → Consider PPO, SAC, or TD3240 - **Discrete** → Consider PPO, A2C, or DQN241 - **MultiDiscrete/MultiBinary** → Use PPO or A2C2422432. **Is sample efficiency critical?**244 - **Yes (expensive simulations)** → Use off-policy: SAC, TD3, DQN, or HER245 - **No (cheap simulations)** → Use on-policy: PPO, A2C2462473. **Do you need fast wall-clock training?**248 - **Yes** → Use PPO or A2C with vectorized environments249 - **No** → Any algorithm works2502514. **Is the task goal-conditioned with sparse rewards?**252 - **Yes** → Use HER with SAC or TD3253 - **No** → Continue with standard algorithms2542555. **Is the environment partially observable?**256 - **Yes** → Use RecurrentPPO257 - **No** → Use standard algorithms258259### Quick Recommendations260261- **Starting out / General tasks:** PPO262- **Continuous control / Robotics:** SAC263- **Discrete actions / Atari:** DQN or PPO264- **Goal-conditioned / Sparse rewards:** SAC + HER265- **Fast prototyping:** A2C266- **Sample efficiency critical:** SAC, TD3, or DQN267- **Partial observability:** RecurrentPPO268269## Training Configuration Tips270271### For On-Policy Algorithms (PPO, A2C)272273```python274# Use vectorized environments for speed275env = make_vec_env(env_id, n_envs=8, vec_env_cls=SubprocVecEnv)276277model = PPO(278 "MlpPolicy",279 env,280 n_steps=2048, # Collect this many steps per environment before update281 batch_size=64,282 n_epochs=10,283 learning_rate=3e-4,284 gamma=0.99,285)286```287288### For Off-Policy Algorithms (SAC, TD3, DQN)289290```python291# Fewer environments, but use gradient_steps=-1 for efficiency292env = make_vec_env(env_id, n_envs=4)293294model = SAC(295 "MlpPolicy",296 env,297 buffer_size=1_000_000,298 learning_starts=10000,299 batch_size=256,300 train_freq=1,301 gradient_steps=-1, # Do 1 gradient step per env step (4 with 4 envs)302 learning_rate=3e-4,303)304```305306## Common Pitfalls3073081. **Using DQN with continuous actions** - DQN only works with discrete actions3092. **Not using vectorized environments with PPO/A2C** - Wastes potential speedup3103. **Using too few environments** - On-policy methods need many samples3114. **Using too large replay buffer** - Can cause memory issues3125. **Not tuning learning rate** - Critical for stable training3136. **Ignoring reward scaling** - Normalize rewards for better learning3147. **Wrong policy type** - Use "CnnPolicy" for images, "MultiInputPolicy" for dict observations315316## Performance Benchmarks317318Approximate expected performance (mean reward) on common benchmarks:319320### Continuous Control (MuJoCo)321- **HalfCheetah-v3**: PPO ~1800, SAC ~12000, TD3 ~9500322- **Hopper-v3**: PPO ~2500, SAC ~3600, TD3 ~3600323- **Walker2d-v3**: PPO ~3000, SAC ~5500, TD3 ~5000324325### Discrete Control (Atari)326- **Breakout**: PPO ~400, DQN ~300327- **Pong**: PPO ~20, DQN ~20328- **Space Invaders**: PPO ~1000, DQN ~800329330*Note: Performance varies significantly with hyperparameters and training time.*331332## Additional Resources333334- **RL Baselines3 Zoo**: Collection of pre-trained agents and hyperparameters: https://github.com/DLR-RM/rl-baselines3-zoo335- **Hyperparameter Tuning**: Use Optuna for systematic tuning336- **Custom Policies**: Extend base policies for custom network architectures337- **Contribution Repo**: SB3-Contrib for experimental algorithms (QR-DQN, TQC, etc.)